flâneur — a map of the web's best reading

Preventing model exfiltration with upload limits

redwoodresearch.substack.com · 4,772 words · saved by 1 readers

Unlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure.

Preventing model exfiltration with upload limits Unlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure. Ryan Greenblatt May 08, 2024 3 Share At some point in the future, AI developers will need to ensure that when they train sufficiently capable models, the weights of these models do not leave the developer’s control. Ensuring that weights are not exfiltrated seems crucial for preventing threat models related to both misalignment and misuse. The challenge of defending model weights has previously been discussed in a RAND report

Explore this link on the map →

related reading