The Privatized AI Cold War

A ModSleuth audit of the private borders inside the U.S.–China AI rivalry.

The first Cold War taught the world to see strategic technology as territory. Knowledge sat inside national borders; states classified it, rationed it, and treated a secret crossing as espionage. The AI rivalry has revived that map at the moment the object on it has ceased to hold still.

On July 22, 2026, the director of the White House Office of Science and Technology Policy said Moonshot AI had distilled Anthropic’s Fable to build K3. Five days later, China’s Ministry of Commerce accused “many US AI firms” of distilling Chinese models and called the American investigations “AI hegemonism.”12 Neither government published evidence.

Policy follows the same geography. Washington has removed the DeepSeek app from intelligence-community systems and briefly ordered Anthropic to suspend foreign-national access to two models.34 Beijing’s service registry contains no American frontier model.5 These are real controls, but they land on applications, providers, procurement, and market access. China’s Qinglang campaign reaches training data, yet never mentions distillation.6

From above, two blocs are taking shape. DeepSeek is Chinese, Anthropic is American, and a crossing between them becomes a national-security event. Below that level, a model’s corporate address and technical lineage come apart.

The second map

ModSleuth draws the other map. It starts from a release, gathers its public reports, cards, code, and linked artifacts, records how each dependency shaped the target, attaches the source evidence, and then follows every dependency upstream.7 The result is not a family tree of checkpoints but a graph of operations: models generating data, filtering corpora, rewriting examples, judging outputs, and steering development.

Across four releases, ModSleuth recovered 1,060 verified dependencies—more than three times the strongest general-purpose baseline—with paths up to eight hops. It found 350 cases of upstream models operating on training data and only 28 of direct weight lineage.7 What a model is “built on” usually sits several documents away from its weights. The code and graph explorer are public.

For this article, I applied that view to recent American and Chinese releases. NVIDIA openly described distilling teachers from DeepSeek, OpenAI, Google, Moonshot, and others.8 Google’s SigLIP initialized the vision encoder in Kimi K2.5; K2.5 outputs later helped bootstrap an American model from Thinking Machines Lab.910 Cursor built Composer 2.5 on Kimi K2.5, while Moonshot trained Kimi-Dev on publicly released trajectories collected with Claude.1112

The paths do not merely cross the border. They loop through it. One longer chain runs from Hangzhou to Shanghai through Santa Clara (Figure 1).

Figure 1One documented relay, Hangzhou to Shanghai by way of Santa Clara
  1. DeepSeek · Hangzhou publishes R1, and every language-model release through the flagship V4-Pro, as open weights under MIT terms.

    Governing text: the models "allow for any modifications and derivative works, including, but not limited to, distillation for training other LLMs."

  2. crosses China → U.S.: no government permission required; the terms travel with the weights

    NVIDIA · Santa Clara generates a 336.6-billion-token corpus of SFT-style training data; DeepSeek and Qwen models produce 314.8 billion of those tokens, or 93.5 percent.

    Governing text: the teachers' licenses; the dataset card warns that trained models "may be subject to redistribution and use requirements in the Qwen License Agreement and the DeepSeek License Agreement."

  3. NVIDIA republishes downstream under its own open terms: Nemotron models under NVIDIA's open model license and, from June 2026, the Linux Foundation's OpenMDW; the Nemotron terminal-task set under CC BY 4.0.

    Governing text: OpenMDW grants rights "to deal in the Model Materials without restriction."

  4. crosses U.S. → China: again with no state instrument in the path

    Shanghai AI Lab trains Intern-S2 with NVIDIA's published Nemotron terminal-task set as the largest single source in its table of executable coding and terminal tasks: 80,000 of 214,854.

    Governing text: NVIDIA's open release terms; the report also adopts Nemotron's warm-up schedule.

ModSleuth reconstructs the chain from public artifacts. Every step was authorized in advance by the text that accompanied the artifact.13141516

No government instrument appears anywhere in the chain.

Another path changes form as it moves. Chinese models generated answers for a Prime Intellect corpus; a later subset kept only difficulty scores derived from their pass rates; Zhipu used those scores to train GLM-5; NVIDIA then trained on 1.3 million GLM-5-generated conversations.171819 Output became metadata, then a model, then new output.

A national label names the company that released the final node. It does not describe the graph inside it. States can block a product or provider; they cannot make these paths resolve into two national supply chains.

The private border

The graph is not borderless. Its borders sit on the edges. For open models, the governing text is the license that travels with the artifact; recursive tracing can reveal obligations inherited through intermediate models and synthetic datasets.7

DeepSeek moved from a V3 license that explicitly defined synthetic-data distillation to plain MIT terms for R1 and every release through V4-Pro.2013 MiniMax moved the other way: its H3 license excludes the United States, the European Union, the United Kingdom, and South Korea, and forbids using outputs to improve another model, while offering a route back through company authorization.212223

The boundary can change inside one product line. Alibaba put Qwen3.8-Max behind a revenue gate, then released a smaller model from the same generation under Apache 2.0 two days later.24 Meta’s Llama naming clause led Ai2 to regenerate a 5.62-billion-token mathematics corpus with Qwen rather than pass the obligation downstream.2526

This is not one line between two countries. It is a set of private borders that shifts by release, model size, revenue, territory, and use. States still control the outer choke points; within the model graph, companies write the rules edge by edge.

The invisible edge

ModSleuth can reconstruct the crossings above because companies documented them. It cannot decide whether Moonshot covertly extracted Fable outputs: that alleged edge appears in no public artifact. This is not a gap that a larger document search will close. It is where public evidence ends and private infrastructure begins.

Open and served models have opposite kinds of visibility. An open license is public, but use of the weights is hard to observe. A served model can be watched closely, but only by its provider. Anthropic blocks commercial Claude access from China through accounts and infrastructure, not under any statute, and only Anthropic holds the complete logs.27

If the accusation is true, the event has national consequences. Yet its operative relationship is Anthropic–Moonshot before it is United States–China: company policy supplies the rule, company infrastructure detects the breach, and company records establish what happened. Anthropic’s quantified report predates K3 by nearly five months and names no accused model; the White House post names K3 but publishes nothing behind the claim.271 K3’s report names only earlier Kimi models as post-training generators.28 The one public experiment that tested K3 called its result inconclusive and could not read Fable’s traces.2930

Settling the matter would require Anthropic logs tied to Moonshot accounts or Moonshot’s training records. Both are private, and an April memorandum routes government evidence back to the affected companies through a private channel.1

That is what makes this more than an AI industry dominated by private firms. Defense contractors also built the weapons of the first Cold War. Here companies perform functions that make the strategic border itself real: they grant passage, patrol it, and keep the record of its breach.

States have not disappeared. They still name the adversary and control chips, procurement, providers, and market access. Their power is broad but coarse. Corporate power is narrower and granular: a company can change the terms of one checkpoint, refuse one account, observe one stream of calls, and retain the only evidence. The proposed BLADE Act would allow intent to be inferred from circumstantial patterns—a legal substitute for direct records held in private.31 No independent channel yet exists for verifying those records, and no public instrument governs distillation at training time.32

The old map has not vanished. Washington and Beijing can harden the perimeter and call the models inside American or Chinese. But beneath the flags, the artifacts keep crossing, returning, and reappearing inside their supposed rivals. ModSleuth can recover the paths left in public; the decisive missing paths remain in corporate systems. This cold war is declared by states. Its inner borders, its patrols, and much of its truth are private.


  1. Michael Kratsios, Director, White House OSTP, post on X, July 22, 2026: “We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model. … To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection.” National Security Technology Memorandum 4, April 23, 2026. — x.com · whitehouse.gov ↩︎ ↩︎ ↩︎

  2. Ministry of Commerce of the PRC, spokesperson Q&A, July 27, 2026: 「据了解,很多美国人工智能企业在研发和训练中蒸馏了中国的模型。」 「这些做法缺乏事实依据,没有法理支撑,在实践上搞双重标准,属于典型的人工智能霸权主义行径。」 — mofcom.gov.cn ↩︎

  3. National Defense Authorization Act for Fiscal Year 2026, Pub. L. 119-60, §6604: “PROHIBITION ON USE OF DEEPSEEK ON INTELLIGENCE COMMUNITY SYSTEMS.” “‘Covered application’ means the DeepSeek application or any successor application or service.” — govinfo.gov ↩︎

  4. Anthropic, “Statement on the US government directive to suspend access to Fable 5 and Mythos 5,” June 12, 2026: “The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees.” Controls lifted June 30; Fable 5 restored July 1. The one customer suit was voluntarily dismissed without a ruling on the merits. — anthropic.com · anthropic.com/redeploying ↩︎

  5. Cyberspace Administration of China, July 10, 2026: 988 generative-AI services filed as of June 30. — cac.gov.cn ↩︎

  6. 人工智能生成合成内容标识办法, published March 14, 2025, in force September 1, 2025: requires labeling of AI-generated content by provider name and content number. 中央网信办 “清朗” campaign, April 30, 2026: polices training-data sourcing compliance; the word 蒸馏 (distillation) does not appear. Phase-one results (July 6, 2026): more than 14,000 AI products handled. — cac.gov.cn/labeling · cac.gov.cn/qinglang ↩︎

  7. Sanjay Adhikesaven, Haoxiang Sun, and Sewon Min, “Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs,” arXiv:2606.12385, June 10, 2026. Across four target releases, ModSleuth recovered 1,060 source-verified dependencies in its unbounded evaluation scope, more than three times the strongest baseline. The paper reports dependency paths up to eight hops; among 1,654 verified forward-reachable relationships, 350 involve upstream operations on training data and 28 involve weight-level lineage. — paper · code · demo ↩︎ ↩︎ ↩︎

  8. nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 model card, August 11, 2026: “During post-training, we generate synthetic data by distilling trajectories, solutions, and translations from strong teacher models.” The generation table lists DeepSeek-V4-Pro, Kimi-K2.5, Qwen variants, GLM-4.6 through GLM-5, MiniMax-M2, gpt-oss-120b, Mixtral, and Gemma. — huggingface.co ↩︎

  9. Kimi Team, “Kimi K2.5,” arXiv:2602.02276, §4.2: “Initialized from SigLIP-SO-400M.” — arxiv.org ↩︎

  10. Thinking Machines Lab, “Inkling,” July 15, 2026: “To bootstrap post-training, we ran an initial SFT on synthetic data generated by open-weights models including Kimi K2.5.” — thinkingmachines.ai ↩︎

  11. Cursor (Anysphere), “Mixture-of-Kittens,” August 4, 2026: “take Kimi 2.5, the base model for Composer 2.5, where H = 7168 and I = 2048.” “MoK delivered an approximately 41% tokens-per-second speedup over our previous DeepEP-based production setup.” — cursor.com ↩︎

  12. Zonghan Yang et al. (Moonshot AI, Tsinghua, Peking University), “Kimi-Dev,” arXiv:2509.23045, §4.1: “5,016 SWE-Agent trajectories collected with Claude 3.7 Sonnet.” — arxiv.org ↩︎

  13. deepseek-ai/DeepSeek-R1 model card, January 20, 2025: “allow for any modifications and derivative works, including, but not limited to, distillation for training other LLMs.” LICENSE file: MIT. DeepSeek-V4-Pro-0813 (August 13, 2026): MIT. — huggingface.co ↩︎ ↩︎

  14. nvidia/Nemotron-Pretraining-SFT-v1 dataset card: a twenty-row per-generator token table stating no total; the 93.5% is the author’s arithmetic. License section: trained models “may be subject to redistribution and use requirements in the Qwen License Agreement and the DeepSeek License Agreement.” — huggingface.co ↩︎

  15. Nemotron 3 Ultra (June 2026) and 3.5 Lightning (August 2026) carry OpenMDW-1.1; Nemotron-Terminal-Synthetic-Tasks carries CC BY 4.0. OpenMDW 1.1: “permission is hereby granted, free of charge, to deal in the Model Materials without restriction.” — openmdw.ai ↩︎

  16. Intern-S2-Preview Team, Shanghai AI Laboratory, arXiv:2608.13505, August 13, 2026. Table 1: Nemotron-Terminal-Synthetic-Tasks 80,000 of 214,854 executable coding and terminal tasks. — arxiv.org ↩︎

  17. PrimeIntellect/SYNTHETIC-2 (Apache-2.0; 353,945 rows): response columns name DeepSeek-R1-0528, Qwen3-32B, and Qwen3-4B as generators. PrimeIntellect/SYNTHETIC-2-RL (155,638 rows): carries no response columns — only prompts, rule-based verifiers, and pass-rate annotations; no license declared. Prime Intellect, Inc. is registered in Delaware. — huggingface.co ↩︎

  18. Zhipu AI, “GLM-5,” arXiv:2602.15763v2, §3.2: competitive-programming RL data “is primarily sourced from Codeforces and representative datasets such as TACO and SYNTHETIC-2-RL.” — arxiv.org ↩︎

  19. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 model card: “Synthetic Chat Reasoning-Off Data from GLM-5 | 646,738” and “Synthetic Chat Reasoning-On Data from GLM-5 | 644,286.” The 1,291,024 total is the author’s sum.huggingface.co ↩︎

  20. DeepSeek-V3 LICENSE-MODEL (repository created December 25, 2024), §1: “‘Derivatives of the Model’ means… distillation methods entailing the use of intermediate data representations or methods based on the generation of synthetic data by the Model for training the other model.” — huggingface.co ↩︎

  21. MiniMax model licenses, October 2025 to August 2026: M2 (“Modified MIT”), M2.5 (split code/model licenses), M2.7 (“NON-COMMERCIAL LICENSE”), M3 (“COMMUNITY LICENSE”), H3 (“COMMUNITY LICENSE AGREEMENT” with Excluded Territories). — huggingface.co/H3 ↩︎

  22. MiniMaxAI/MiniMax-H3 LICENSE, definition 5: “‘Excluded Territories’ means the European Union, the United Kingdom, the Republic of Korea and the United States of America.” §V.3: “You may not use the MiniMax H3 Works or any of their Outputs or results to improve any other artificial intelligence model.” — huggingface.co ↩︎

  23. MiniMax, X post, August 4, 2026: “MiniMax H3 can be licensed for deployment in the US, EU, UK, and South Korea through MiniMax’s formal authorization process.” License Q&A: “MiniMax is also involved in ongoing copyright-related legal proceedings specifically concerning generative video AI.” — x.com · huggingface.co ↩︎

  24. Qwen/Qwen3.8-2.4T-A95B LICENSE (“Qwen3.8-Max License”, August 12, 2026): separate agreement required above $50M yearly revenue. Qwen/Qwen3.8-27B (August 14, 2026): Apache-2.0. — huggingface.co/Max · huggingface.co/27B ↩︎

  25. Meta, “Llama 4 Community License Agreement,” §1.b.i: “you shall also include ‘Llama’ at the beginning of any such AI model name.” §2: above 700M monthly active users, “you must request a license from Meta, which Meta may grant to you in its sole discretion.” — github.com ↩︎

  26. Team Olmo (Ai2), “Olmo 3,” arXiv:2512.13961, appendix: “SwallowMath… was rewritten using a Llama model, which would require that any model trained on this data would need to have ‘Llama’ in the name… To provide truly open data, we mirror the generation of this dataset, but use Qwen3 32B… This yields a 5.62B token dataset we refer to as CraneMath.” — arxiv.org ↩︎

  27. Anthropic, “Detecting and preventing distillation attacks,” February 23, 2026: “over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts.” Per lab: MiniMax over 13 million, Moonshot over 3.4 million, DeepSeek over 150,000. “Distillation is a widely used and legitimate training method.” “For national security reasons, Anthropic does not currently offer commercial access to Claude in China.” — anthropic.com ↩︎ ↩︎

  28. Kimi Team (Moonshot AI), “Kimi K3,” arXiv:2607.24653. §2.4: “A key departure from Kimi K2.5 is that we train Kimi K3 vision encoder, MoonViT-V2, entirely from scratch… We depart from this practice primarily for training stability.” §4.1.1: “we synthesize data trajectories using domain-specialized models from the prior Kimi series.” — arxiv.org ↩︎

  29. Alexander Panfilov et al., “Stealing Reasoning Traces from Proprietary LLM APIs,” arXiv:2608.09867, August 10, 2026. “After Bonferroni correction, only Kimi-K3 mean differences remain significant.” “Sol and Opus prefills change five out of six models.” Authors’ verdict: “suggestive but inconclusive… cannot establish a causal claim of memorization or distillation.” — arxiv.org ↩︎

  30. Panfilov et al., arXiv:2608.09867, Table 1 caption: “the thinking traces of any model can be replayed by any other, except Fable 5’s thoughts.” — arxiv.org ↩︎

  31. BLADE Act, S. 5252, 119th Congress, introduced August 5, 2026. §3(7)(B): “the purpose of extraction may be inferred from the totality of circumstances, including– (i) the volume, structure, pattern, coordination, or timing of the extraction activity; (ii) the concentration of extractions on specific model capabilities; (iii) the use of multiple accounts in a coordinated manner; or (iv) the correlation of extraction activity within the development timeline of another artificial intelligence model.” — govinfo.gov ↩︎

  32. Executive Order 14409, signed June 2, 2026. §3(c): “Nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models.” The framework was delivered on its August 1 deadline and not published; Axios reported it defines covered models as closed-source only. — federalregister.gov · axios.com ↩︎