Matrix-based Filtering and Load-balancing Algorithm for Efficient Similarity Join Query Processing in Distributed Computing Environment Yang, Hyeon-Sik; Jang, Miyoung; Chang, Jae-Woo;
As distributed computing platforms like Hadoop MapReduce have been developed, it is necessary to perform the conventional query processing techniques, which have been executed in a single computing machine, in distributed computing environments efficiently. Especially, studies on similarity join query processing in distributed computing environments have been done where similarity join means retrieving all data pairs with high similarity between given two data sets. But the existing similarity join query processing schemes for distributed computing environments have a problem of skewed computing load balance between clusters because they consider only the data transmission cost. In this paper, we propose Matrix-based Load-balancing Algorithm for efficient similarity join query processing in distributed computing environment. In order to uniform load balancing of clusters, the proposed algorithm estimates expected computing cost by using matrix and generates partitions based on the estimated cost. In addition, it can reduce computing loads by filtering out data which are not used in query processing in clusters. Finally, it is shown from our performance evaluation that the proposed algorithm is better on query processing performance than the existing one.
Konstantin Shvachko, Hairong Kuang, Sanjay Radia, Robert Chansler "The hadoop distributed file system," Mass Storage Systems and Technologies (MSST), pp.1-10, 2010.
Jeffrey Dean and Sanjay Ghemawat, "MapReduce: simplified data processing on large clusters," Communications of the ACM, Vol.51, Issue.1, pp.107-113, 2010.
Surajit Chaudhuri, Venkatesh Ganti, and Raghav Kaushik, "A primitive operator for similarity joins in data cleaning," Data Engineering, p.5, 2006.
A. Metwally, D. Agrawal, and A. El Abbadi, "DETECTIVES: DETEcting Coalition hiT Inflation attacks in adVertising nEtworks Streams," Proceedings of the 16th WWW International Conference on World Wide Web, pp.241-250, 2007.
A. Z. Broder, S. C. Glassman, M. S. Manasse, and G. Zweig, "Syntactic clustering of the web," Computer Networks, pp.1157-1166, 1997.
T. C. Hoad and J. Zobel, "Methods for identifying versioned and plagiarized documents," JASIST, Vol.54, Issue.3, pp.203-215, 2003.
Yasin N. Silva and Jason M. Reed, "Exploiting MapReduce-based similarity joins," Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data, pp.693-696, 2012.
Ahmed Metwally and Christos Faloutsos, "V-smart-join: A scalable mapreduce framework for all-pair similarity joins of multisets and vectors," Proceedings of the VLDB Endowment, Vol.5, No.8, pp.704-715, 2012.
Alper Okcan and Mirek Riedewald, "Processing theta-joins using MapReduce," Proceedings of the 2011 ACM SIGMOD International Conference on Management of data ACM, pp.949-960, 2011.