flâneur

Joshua Achiam on X: "A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do" / X

x.com · 226 words · saved by 1 readers

A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do not think it makes …

A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do not think it makes sense to elevate as a principle the idea that the chain of thought must remain legible to humans. I would go so far as to say that strategies predicated on that principle are definitely doomed, in that they will not work eventually, and we should not depend on them or take enduring reassurance from…

saved by

related reading