Small edits, large models: How Wikipedia advocacy shapes LLM values | Zenodo
Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in nearly every major language model training dataset and is weighted more heavily than web-crawled text. The Pro-Animal Wikipedians (PAW), a group of advocates who add sourced animal welfare content to relevant articles, have made 125 edits across 115 pages. Using gradient-based data attribution (Bergson; Lucia and Belrose 2026; MAGIC; Ilyas and Engstrom 2025), we traced how these edits inuence language model behavior. TrackStar retrieval attribution on Llama 3.1 8B found that PAW-edited sections made up 68% of the highest-attributed documents for animal welfare queries (p < 0.0001) but only 52% for unrelated queries about the same companies (p = 0.53): the model links PAW content specically to animal welfare topics, not to the entities in general. MAGIC counterfactual inuence estimation on Llama-3.2-1B (validated at ρ = 0.810.95, p ≤ 0.005) confirmed that PAW content causally shifts the model's animal welfare predictions more than control content does. When we finetuned separate models on PAW content versus control content, each model performed better speciccally on the type of text it was trained on, with no crossover beneit: the PAW-trained model cut perplexity on animal welfare text from 12.4 to 8.4, while the control-trained model cut perplexity on control text from 16.1 to 11.4. A small, coordinated Wikipedia editing campaign therefore measurably shapes how language models handle the topics those edits address, making Wikipedia editing a practical, low-cost way for advocacy organizations to inuence AI systems.