Poisoning Attacks on LLMs: How Adversaries Corrupt Models and What Defenders Can Do
Poisoning attacks corrupt the data, weights, or fine-tuning process of large language models. This post breaks down the attack surface, walks through recent incidents, and lays out the defenses working teams can ship today.