flâneur

book3n.dvi

infolab.stanford.edu · 10,194 words · saved by 1 readers

N/A

72 Chapter 3 Finding Similar Items A fundamental data-mining problem is to examine data for “similar” items. We shall take up applications in Section 3.1, but an example would be looking at a collection of Web pages and finding near-duplicate pages. These pages could be plagiarisms, for example, or they could be mirrors that have almost the same content but differ in information about the host and about other mirrors. The naive approach to finding pairs of similar items requires us to look at ev- ery pair of items. When we are dealing with a large dataset, looking at all pairs of…

related reading