✳flâneur — a map of the web's best reading
The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- GitHub - mc2-project/mc2: A Platform for Secure Analytics and Machine Learning · GitHubgithub.com
- Lapis Labslapis.rocks
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness · GitHubgithub.com
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scaling · GitHubgithub.com
- GitHub - SoyGema/pulling_ace · GitHubgithub.com
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- BIG-bench/bigbench/benchmark_tasks/keywords_to_tasks.md at main · google/BIG-bench · GitHubgithub.com
- GitHub - open-edge-platform/anomalib: An anomaly detection library comprising state-of-the-art algorithms and features such as experiment management, hyper-parameter optimization, and edge inference. · GitHubgithub.com
- GitHub - suzana-ilic/study_model_behavior: Model Behavior Study Group · GitHubgithub.com
- MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity - METRmetr.org
- GitHub - adamlyttleapps/notchy · GitHubgithub.com
- Making sure you're not a bot!inria.hal.science