Designing and Interpreting Probes · John Hewitt
Human languages are wild, delightfully tricky phenomena, exhibiting structure and variation, suggesting general rules and then breaking them. In some part, it’s for this reason that the best machine learning methods for modeling natural language are extremely flexible, left to their own devices to build internal representations of sentences. However, this flexibility comes at a cost. The most popular models in natural language processing research at this time are considered black boxes, offering few explicit cues to what they learn about how language works. Earlier machine learning methods for NLP learned combinations of linguistically motivated features—word classes like noun and verb, syntax trees for understanding how phrases combine, semantic labels for understanding the roles of entities—to implement applications involving understanding some aspects of natural language. Though it was difficult at times to understand exactly how these combinations of features led to the decisions m
Explore this link on the map →