ScruTool
News

White House Blames Distillation for Kimi K3's Strength. Researchers Say the Timeline Rules It Out

White House accuses Moonshot AI of distilling Anthropic's Fable into Kimi K3. AI researchers question the evidence, timeline, and technical feasibility.

Jul 23, 2026 9 min read

Kratsios and Bessent accuse Moonshot of copying Anthropic's Fable to build the largest open-weight model yet. The people who study how these systems are trained are not buying the strongest version of the claim.

WASHINGTON — The Trump administration has accused Chinese startup Moonshot AI of covertly copying Anthropic's Fable model to build Kimi K3, and Treasury Secretary Scott Bessent has raised the prospect of sanctions. AI researchers who study how these systems are actually built say the timeline behind the accusation does not hold up.

Michael Kratsios, director of the White House Office of Science and Technology Policy, laid out the charge in a post on X on Wednesday. He wrote that his office has information indicating Moonshot distilled Fable for the development of K3, and that the company built an internal platform to run large-scale distillation against U.S. models while switching between access methods to avoid detection. He also alleged that Moonshot had obtained servers fitted with Nvidia's GB300 chips and accessed GB300s in Thailand, hardware that is banned from export to China.

Kratsios did not share the evidence behind the claim. Moonshot has not responded to the accusation.

The accusation, and the escalation behind it

Bessent set the tone a day earlier. In a Fox Business interview on Tuesday, he said the government had detected watermarks from U.S. large language models inside many Chinese systems and called it unacceptable, adding that the administration would examine the issue over the coming days or weeks. He also floated a second idea: that American companies running Chinese models might have to disclose that fact to their customers.

By Wednesday, Bessent had sharpened the threat. "Open source is not open season on American IP," he wrote on X, warning that covert, industrial-scale distillation attacks crossing into intellectual property theft would put sanctions and Entity List designations on the table.

Distillation is the point of contention. The technique feeds the outputs of a stronger model to a weaker one, training the weaker model to imitate the stronger. Done openly and at small scale, it is legal and ordinary across the industry. The administration's objection is to a covert, industrial version of it aimed at lifting proprietary capability wholesale.

This is not the first time Anthropic's models have been at the center of such a claim. Earlier this year the company alleged that Moonshot generated more than 3.4 million Claude exchanges through fraudulent accounts to pull out capabilities spanning reasoning, coding, tool use, and computer vision. Anthropic said the metadata tied the activity to senior Moonshot employees and described it as deliberate capability extraction rather than ordinary use. Beijing called the earlier round of allegations groundless.

Why researchers doubt the distillation story

The gap in this week's accusation is between what officials are asserting and what technical researchers will stand behind.

Fable has been publicly available only since July 1. Moonshot released Kimi K3 on July 16. That leaves roughly two weeks.

"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch. He argued there simply was not enough time to query a model at that scale, generate the data, train on it, and ship a frontier system in two weeks.

Nathan Lambert, an AI researcher at the Allen Institute for AI, made a broader point on a podcast released this week. As Chinese models close on the frontier and training shifts toward reinforcement learning, he argued, distillation matters less than it once did. If copying outputs alone could transfer frontier capability, every lab would already have caught up to models like Kimi K3 by doing the same thing. That has not happened.

There is a narrower version of the claim that researchers do find plausible. Supervised fine-tuning, where a model trains on the prompts and responses of another, is where a model "picks up its manners," in Lambert's phrase. It is also what can leave a Chinese model occasionally identifying itself as Claude. Those surface fingerprints are the likeliest source of the watermarks Bessent referenced. They point to contact with Anthropic's outputs. They do not explain frontier reasoning.

Reproducing capability at Fable's level would require reinforcement learning, not fine-tuning. In many setups that means running an agent of the larger model to grade the smaller one's answers and adjusting from the score. Doing that through a frontier lab's API would be, in the researchers' description, prohibitively expensive, slow enough to become a bottleneck, and quite possibly no help to performance at all. Large reinforcement learning runs can involve tens of millions of agents.

Hancock also pushed back on the framing that Chinese labs are merely riding on stolen American work. One of Moonshot's founders was a Carnegie Mellon PhD student, he noted, and the teams are staffed with capable researchers. If American models stopped advancing tomorrow, he said, China's progress would slow but continue.

Distillation is not unique to Chinese labs either. Elon Musk testified earlier this year that his AI company distilled OpenAI's models while developing Grok and described the practice as common. The boundary between distillation and building synthetic training datasets is often blurry.

What Kimi K3 actually is

Stripping away the accusation, the model at the center of it is a genuine milestone.

Moonshot released Kimi K3 on July 16 as a 2.8-trillion-parameter Mixture-of-Experts model that activates only 16 of its 896 experts per token, keeping compute cost far below what a dense model that size would demand. It reads a one-million-token context window and accepts text, images, and video. Moonshot calls it the largest open-weight model built to date. The company has promised to publish the full weights by July 27.

The benchmarks are what unsettled U.S. labs. Kimi K3 debuted at number one on Arena's Frontend Code leaderboard with 1,679 points, edging past Claude Fable 5 at 1,631 and GPT-5.6 Sol at 1,618. That was a 17-place jump from the previous Kimi flagship. On the broader Artificial Analysis Intelligence Index it scored 57 and ranked fourth out of 189 models, level with Claude Opus 4.8 and GPT-5.5, trailing only Fable 5 and GPT-5.6 Sol.

The catch for anyone imagining a homegrown copy running on a laptop: the weights come to roughly 1.4 terabytes even in their native 4-bit format, and Moonshot's own deployment guidance calls for 64 or more accelerators. This is a datacenter model.

The chip question is separate, and murkier

Kratsios bundled a second charge into the same post: that Moonshot obtained GB300-equipped servers and reached GB300s sitting in Thailand. Those chips belong to Nvidia's Blackwell generation and cannot legally be sold to Chinese companies.

Whether Moonshot copied Fable and whether it trained on restricted hardware are two different questions, and the second does not answer the first. Compute access is about how the model was trained at scale, not about whose outputs went into it.

A black market for banned chips does exist. Sam Bresnick, a research fellow at Georgetown's Center for Security and Emerging Technology, has pointed to it, and in May the founder of U.S. server builder Supermicro was indicted for smuggling advanced chips into China. Bresnick has argued for know-your-customer rules covering data centers worldwide, so that any company running a large training run on state-of-the-art hardware has to be identified along with what it is doing.

The Biden administration's Commerce Department proposed federal know-your-customer rules for data centers in 2024. No further progress appears to have been made under President Trump. Exporters shipping advanced chips abroad are still supposed to ensure the hardware is used only for approved purposes.

The real fight is over open weights

The accusation landed at a loaded moment. According to Axios reporting, the administration has been weighing how to curb Chinese open-weight models in the U.S., and the release of Kimi K3 has revived that push after earlier efforts stalled.

Officials have considered several routes since 2025, including adding Chinese labs to the Entity List, issuing security advisories, and an executive order that would hold U.S. companies liable for breaches tied to hosting Chinese models. Those plans were shelved over worries that they would choke domestic innovation. Personnel changes shifted the balance. Sriram Krishnan, a White House adviser who opposed intervention, has left, and voices favoring tighter controls now carry more weight. A source told Axios the government may not need a formal ban at all, describing a slower, more durable playbook built on procurement rules, Entity List threats, and public pressure aimed at U.S. firms that use the models.

An outright ban would be hard to enforce in any case. Once open weights are downloaded, they cannot be recalled or switched off remotely.

That is exactly where critics see the accusation doing double duty. David Sacks, an outside White House AI adviser, wrote on X that leading closed labs, already a duopoly in model revenue, want the government to eliminate their open-source competition. Market analysts have leveled the same charge of regulatory capture. Anthropic and OpenAI executives have argued the opposite case: Anthropic CEO Dario Amodei has warned that freely downloadable models with advanced cyber capabilities could cause serious harm, and OpenAI's Dean Ball has argued that near-free AI systems would starve frontier development of funding.

Representative Ted Lieu of California pointed to the contradiction running underneath all of it. The administration is threatening sanctions over alleged AI intellectual property theft while continuing to clear sales of high-performance AI chips to China.

Community

Discussion

Join the discussion and share your perspective.

Related Articles