Skip to main content

Command Palette

Search for a command to run...

The Distillation Dilemma: Anthropic, OpenAI, and the Security Predicament for Frontier AI Labs

AI Series: Mid-Series Special 12

Updated
•6 min read•View as Markdown
The Distillation Dilemma: Anthropic, OpenAI, and the Security Predicament for Frontier AI Labs
R

I am a software engineer.

Hello 👽,

This is going to be maybe the first back-to-back mid-series special, enjoy it while it lasts 😁.

Let's do what we've come here to do 🕶️

A few weeks ago, Anthropic made a post on Twitter (I refuse to call it X) and published an article titled Detecting and preventing distillation attacks on their news blog. You can read about it directly, as I have attached the Twitter and news blog links.

But let me summarise it for you. Anthropic claimed that some companies, notably DeepSeek, Moonshot AI, and MiniMax, used model distillation attacks to extract data to train and improve their models using their APIs for a fraction of the cost. So first of all, what is model distillation?

Model Distillation

Model distillation (also known as knowledge distillation) is a technique in machine learning where a large, powerful model is used to train a smaller, faster model so that the smaller model can mimic the behavior of the larger one.

The goal is to compress the knowledge of a big model into a smaller model without losing too much performance.

Here's how it works:

  1. Frontier Labs like Anthropic, OpenAI, and Google Gemini train their initial base models (foundation models) and fine-tune them across different fields to have general knowledge, which is immense work. In our previous series, Chapter 34: BERT in its Element, we learned about model pre-training and fine-tuning. You can check it out to understand the difference. The pre-training step is for the model to understand general language and context mainly. The fine-tuning step is where other specific knowledge fields are trained. Fields like programming, physics, medicine, poetry, and so on. The fine-tuning step is the most expensive, because the frontier labs would need to use real professionals to give them accurate data, and it also takes more time. The pre-training step is not cheap either, as a lab would need thousands of the best NVIDIA GPUs to pre-train a model. After this step, they get large base models like GPT-4, GEMINI, and BERT.

  2. Model distillation happens after getting the base models. This is to create smaller, cheaper models for their customers. What happens here:

    1. The labs take a smaller pre-trained model that has not been fine-tuned to train through model distillation. They then generate hundreds of thousands to millions of questions and ask the base model.

    2. The fine-tuned base model returns not only the answer but the probability distribution of each answer. These are known as soft labels, or synthetic data. So, for example, if the question is "Generate an image of a cat", the answer and probability distribution would likely look like "Cat: 0.90, Dog: 0.07, Fox: 0.03". This basically means that the feature parameters of the model would reduce drastically, which would make it faster and smaller compared to the fine-tuned large base model. This is how smaller and lighter models like GPT-5 Mini, Claude 3.5 Haiku, and Gemini Flash were created.

Why are distilled models smaller and faster?

Here, the frontier labs use a base model with a small number of parameters, like Llama-8B, which has 8Billion parameters compared to a full-fledged model like Llama-405B, which has 405Billion parameters. Then, using model distillation, fine-tune Llama-8B from Llama-405B's answers and probability distribution, using techniques like the rubric-based grading tasks that function as a reward model for reinforcement learning. So at the end of the day, the Llama-405B is "distilled" into Llama-8B, and having fewer parameters, it is indeed going to be smaller and faster.

A very good analogy to this is an experienced veteran who writes down not only his life lessons but also answers thousands of questions about different situations. A novice studies those answers and learns in months what took the veteran a lifetime to learn.

Now that we know about Model Distillation, why is Anthropic complaining? This is because of what's known as distillation attacks.

Distillation Attacks

A distillation attack (or model-extraction attack) occurs when an adversary systematically queries a proprietary LLM via its API to harvest millions of input-output pairs. This "synthetic textbook" is then used to train a competing model that mimics the original’s reasoning and capabilities at a fraction of the original R&D cost.

Anthropic and OpenAI raised an alarm on industrial-scale distillation attacks on their platforms. They've set up tools and systems to try to curtail these attacks with Product, API, and model-level safeguards.

However, this issue is likely going to simply get worse.

Why I think distillation attacks would only get worse

This is what Anthropic mentioned in their blog

For national security reasons, Anthropic does not currently offer commercial access to Claude in China, or to subsidiaries of their companies located outside of the country.

To circumvent this, labs use commercial proxy services which resell access to Claude and other frontier AI models at scale. These services run what we call “hydra cluster” architectures: sprawling networks of fraudulent accounts that distribute traffic across our API as well as third-party cloud platforms. The breadth of these networks means that there are no single points of failure. When one account is banned, a new one takes its place. In one case, a single proxy network managed more than 20,000 fraudulent accounts simultaneously, mixing distillation traffic with unrelated customer requests to make detection harder.

Several notable personalities, including Elon Musk, think this is "double speak" (if you get my 1984 drift). Because Anthropic itself and other frontier labs trained their models on data from the internet without permission from us all.

But the key point here is that America literally stopped selling NVIDIA GPUs to China and other countries, limiting their ability to pre-train LLMs, as NVIDIA GPUs are light-years ahead of every other alternative in the market. You can't just get an NVIDIA GPU outside America, as it is extremely difficult.

Companies outside America also have the capital, but since they don't have access to NVIDIA GPUs, they're forced to be creative. As the saying goes, "necessity is the mother of invention".

We get to thank Meta for creating open-source llama models for us. I wrote a piece about this sometime back: DeepSeek, Meta, and the rise of open-source LLM models. Without open-source models, none of us outside the big tech companies in the US would be able to access state-of-the-art LLMs.

How long can America keep gatekeeping this technology, and is it worth it? Time will tell.

Hope you enjoyed this piece. Let me know what you think about all these 💭

The end.

See you in the next one. This time around, it has to be BERT implementation lol, I keep posting it 😎

⬅️ Previous Chapter - Research: The Best AI Agents Failed at 96% of Jobs