flâneur

Research | AI Alignment Foundation

aialignmentfoundation.org · 181 words · saved by 1 readers

What we're funding and accelerating to solve alignment.

What we're funding and accelerating to solve alignment. Read more about Modular Pretraining Enables Access Control Modular Pretraining Enables Access Control Dual-use knowledge enables models to assist us with the most difficult and demanding tasks in science, but it also empowers people who would use that knowledge to cause harm. Pre-training with GRAM enables knowledge to be siloed and turned on or off when deployed, so that a single model can be both safe and powerful. Read more about Self-Interpretation in Language Models via Adapter Probes Self-Interpretation in Language Models via…

saved by

related reading