Keep pulling the thread on Dan Hendrycks.
The compressibility of model weights significantly heightens the exfiltration risk for large language models (LLMs).
Attackers can achieve 16x to 100x compression on large language model weights with minimal trade-offs by tailoring compression for exfiltration and relaxing decompression constraints.
Compressing large language model weights by 16x to 100x can reduce the time for an attacker to illicitly transmit them from a server from months to days.
Forensic watermarking is an effective and cheap defense for mitigating the risk of model weight exfiltration.