Downloads: 122 | Views: 280
Research Paper | Computer Science & Engineering | India | Volume 2 Issue 11, November 2013 | Popularity: 6.2 / 10
Exploration of Data Mining Techniques in Record Deduplication
R. Gayathri, A. Malathi
Abstract: In todays business world, the database plays a vital role in decision making. As the organization grows, the size of the database also gets increased. This enormous growth in the database size leads to a problem of dirty data. Dirty data is the replicated data in the database which causes some issues like performance degradation, increasing operational cost and the lack of quality. This can be removed by the process of record deduplication. The record deduplication refers to identifying the same entity with different representations. Further cleaning and removing of replica in the repository become a mandatory work. Thus this paper surveys some of the record deduplication approaches. Also it compares with three approaches to record deduplication such as genetic programming, Modified BAT algorithm, and firefly algorithm approach with its limitation and advantages on all the three got discussed.
Keywords: Record Deduplication, preprocessing, Cleaning, Dirty data, genetic programming, mbat algorithm, firefly algorithm
Edition: Volume 2 Issue 11, November 2013
Pages: 216 - 219
Make Sure to Disable the Pop-Up Blocker of Web Browser
Similar Articles
Downloads: 2 | Weekly Hits: ⮙2 | Monthly Hits: ⮙2
Informative Article, Computer Science & Engineering, India, Volume 12 Issue 2, February 2023
Pages: 1106 - 1106Cyber Warfare by Chinese hackers: The AIIMS Story
Dr. Yatu Rani, Sarthak Jain
Downloads: 4 | Weekly Hits: ⮙1 | Monthly Hits: ⮙4
Research Paper, Computer Science & Engineering, United States of America, Volume 13 Issue 10, October 2024
Pages: 2042 - 2049Intelligent Sentiment Prediction in Social Networks leveraging Big Data Analytics with Deep Learning
Maria Anurag Reddy Basani
Downloads: 106 | Weekly Hits: ⮙1 | Monthly Hits: ⮙1
Survey Paper, Computer Science & Engineering, India, Volume 3 Issue 11, November 2014
Pages: 1850 - 1856A Review on Detection of Outliers Over High Dimensional Streaming Data Using Cluster Based Hybrid Approach
Abhishek B. Mankar, Namrata Ghuse
Downloads: 108
M.Tech / M.E / PhD Thesis, Computer Science & Engineering, India, Volume 3 Issue 11, November 2014
Pages: 1431 - 1434Robust Phase-Based Binarization Model to Improve Degraded Document Images
Samreen M. Shaikh A., Mayur S. Burange
Downloads: 109
Review Papers, Computer Science & Engineering, India, Volume 4 Issue 1, January 2015
Pages: 2180 - 2182Review of Improved Cross Redundant Data Cleaning Algorithm for RFID and WSN Integration
Jayashri M. Dupare, N. U. Sambhe