flâneur

Resources on the AI risk landscape – Ben Pomeranz

benpomeranz.com · 1,322 words · saved by 2 readers

Written for some friends in < 2 <2 hours, with the goal of giving them some decent sense of the "AI risk landscape." This is by no means a definitive list. Thanks to mr. Cleo Nardo for providing some inspiration here. Many sources are taken from this AI futurism reading list, which goes much more in depth. If there is a subject that appears to be important that is missing, which there certainly is, please ask me. Exercises optional. This falls under the broader category of Mechanistic Interpretability, or Mech interp. Mech interp is very cool, but I unfortunately don’t think it will be “solved” before we get superhuman AI researchers. On the other hand, without “solving” it we can still get a lot short to medium term safety gains by reading AIs minds. There are a lot of cool tools to do this. I include this section because I think it provokes a lot of good questions, but in the end I believe that our current ability to do an ok job of reading AI minds is less important than you’d naiv

Written for some friends in <2<2 hours, with the goal of giving them some decent sense of the "AI risk landscape." This is by no means a definitive list. Thanks to mr. Cleo Nardo for providing some inspiration here. Many sources are taken from this AI futurism reading list, which goes much more in depth. If there is a subject that appears to be important that is missing, which there certainly is, please ask me. Exercises optional. Data, trends, and evals: A+: METR time horizons (AKA “The METR Graph”): The most important graph in AI capabilities central to tons of forecasting Skim…

saved by

related reading