The Four Generations of Entity Resolution
暫譯: 實體解析的四個世代
Papadakis, George, Ioannou, Ekaterini, Thanos, Emanouil
- 出版商: Morgan & Claypool
- 出版日期: 2021-03-16
- 售價: $2,250
- 貴賓價: 9.5 折 $2,138
- 語言: 英文
- 頁數: 172
- 裝訂: Quality Paper - also called trade paper
- ISBN: 1636390560
- ISBN-13: 9781636390567
海外代購書籍(需單獨結帳)
商品描述
Entity Resolution (ER) lies at the core of data integration and cleaning and, thus, a bulk of the research examines ways for improving its effectiveness and time efficiency. The initial ER methods primarily target Veracity in the context of structured (relational) data that are described by a schema of well-known quality and meaning. To achieve high effectiveness, they leverage schema, expert, and/or external knowledge. Part of these methods are extended to address Volume, processing large datasets through multi-core or massive parallelization approaches, such as the MapReduce paradigm. However, these early schema-based approaches are inapplicable to Web Data, which abound in voluminous, noisy, semi-structured, and highly heterogeneous information. To address the additional challenge of Variety, recent works on ER adopt a novel, loosely schema-aware functionality that emphasizes scalability and robustness to noise. Another line of present research focuses on the additional challenge of Velocity, aiming to process data collections of a continuously increasing volume. The latest works, though, take advantage of the significant breakthroughs in Deep Learning and Crowdsourcing, incorporating external knowledge to enhance the existing words to a significant extent.
This synthesis lecture organizes ER methods into four generations based on the challenges posed by these four Vs. For each generation, we outline the corresponding ER workflow, discuss the state-of-the-art methods per workflow step, and present current research directions. The discussion of these methods takes into account a historical perspective, explaining the evolution of the methods over time along with their similarities and differences. The lecture also discusses the available ER tools and benchmark datasets that allow expert as well as novice users to make use of the available solutions.
商品描述(中文翻譯)
實體解析(Entity Resolution, ER)是數據整合和清理的核心,因此,大部分研究都在探討如何提高其有效性和時間效率。最初的 ER 方法主要針對結構化(關聯)數據中的真實性(Veracity),這些數據由已知質量和意義的模式(schema)描述。為了實現高效能,它們利用模式、專家和/或外部知識。其中一些方法被擴展以應對數據量(Volume),通過多核心或大規模並行化方法處理大型數據集,例如 MapReduce 範式。然而,這些早期基於模式的方法不適用於網絡數據,因為網絡數據充斥著大量、雜訊、半結構化和高度異質的信息。為了應對多樣性(Variety)的額外挑戰,最近的 ER 研究採用了新穎的、鬆散的模式感知功能,強調可擴展性和對雜訊的穩健性。另一條當前研究的方向則專注於速度(Velocity)的額外挑戰,旨在處理不斷增長的數據集。最新的研究則利用深度學習(Deep Learning)和眾包(Crowdsourcing)方面的重大突破,結合外部知識顯著增強現有方法。
本次綜合講座將 ER 方法根據這四個 V 所帶來的挑戰劃分為四個世代。對於每一個世代,我們概述相應的 ER 工作流程,討論每個工作流程步驟的最新方法,並呈現當前的研究方向。對這些方法的討論考慮了歷史視角,解釋了這些方法隨著時間的演變及其相似性和差異性。講座還討論了可用的 ER 工具和基準數據集,這些工具和數據集使專家和新手用戶都能利用現有的解決方案。
作者簡介
George Papadakis, National and Kapodistrian University of Athens
George Papadakis is a research fellow at the National and Kapodistrian University of Athens, Greece. He also worked at the NCSR "Demokritos," National Technical University of Athens (NTUA), L3S Research Center, and “Athena” Research Center. He holds a Ph.D. in Computer Science from the University of Hanover and a Diploma in Electrical Computer Engineering from NTUA. His research interest focuses on web data mining.
Ekaterini Ioannou, Tilburg University
Ekaterini Ioannou is an Assistant Professor at Tilburg University, the Netherlands. Prior, she worked as an Assistant Professor at Eindhoven University of Technology, as a Lecturer at the Open University of Cyprus, an adjunct faculty at EPFL in Switzerland, a research collaborator at the Technical University of Crete, and as an Independent Expert for the European Commission. Her research focuses on information integration with an emphasis on the challenges of man aging data with uncertainties, heterogeneity or correlations, and, more recently, on achieving a deeper integration of information extraction tasks within databases, and on efficiently retrieving analytics over graphs/hypergraphs with evolving data.
Emanouil Thanos, Katholieke Universiteit Leuven
Emanouil Thanos is a Ph.D. candidate at CODeS research group of KU Leuven, under the supervision of Prof. Greet Vanden Berghe. He holds a Diploma in Electrical and Computer Engineering from the National Technical University of Athens and a joint Master in Com putational Logic from TU Dresden, FU Bolzano, and UN Lisbon. He has also worked as a research associate at National ICT Australia and the University of Athens. His research inter ests focus on combinatorial optimization and operations research.
Themis Palpanas, University of Paris, France & French University Institute
Themis Palpanas is Senior Member of the French University Institute (IUF), and Professor of Computer Science at the University of Paris (France) where he is director of the Data Intel ligence Institute of Paris (diiP), and of the Data Intensive and Knowledge Oriented Systems (diNo) group. He is the author of two French patents and nine U.S. patents, three of which have been implemented in world-leading commercial data management products. He is the recipient of three Best Paper awards and the IBM Shared University Research (SUR) Award. He is currently serving in the Board of Trustees for the Very Large Data Bases (VLDB) Endowment, as Editor in Chief for BDR Journal, Editorial Advisory Board member for IS Journal, and in the Senior Program Committee of SIGMOD 2021.
作者簡介(中文翻譯)
喬治·帕帕達基斯,雅典國立與卡波迪斯特里亞大學
喬治·帕帕達基斯是希臘雅典國立與卡波迪斯特里亞大學的研究員。他曾在國家科學研究中心「德莫克里托斯」、雅典國立技術大學 (NTUA)、L3S 研究中心以及「雅典」研究中心工作。他擁有漢諾威大學的計算機科學博士學位,以及雅典國立技術大學的電氣計算機工程學文憑。他的研究興趣集中在網頁數據挖掘上。
艾卡特里尼·伊奧安努,蒂爾堡大學
艾卡特里尼·伊奧安努是荷蘭蒂爾堡大學的助理教授。在此之前,她曾擔任埃因霍溫科技大學的助理教授、塞浦路斯開放大學的講師、瑞士洛桑聯邦理工學院的兼職教員、克里特技術大學的研究合作者,以及歐洲委員會的獨立專家。她的研究專注於信息整合,特別是處理不確定性、異質性或相關性數據的挑戰,最近則著重於在數據庫中實現信息提取任務的更深層整合,以及高效檢索隨著數據演變的圖形/超圖的分析。
埃馬努伊爾·塔諾斯,魯汀大學
埃馬努伊爾·塔諾斯是魯汀大學 CODeS 研究小組的博士候選人,指導教授為 Greet Vanden Berghe。他擁有雅典國立技術大學的電氣與計算機工程文憑,以及德累斯頓工業大學、博爾扎諾大學和里斯本大學的計算邏輯聯合碩士學位。他還曾在國家 ICT 澳大利亞和雅典大學擔任研究助理。他的研究興趣集中在組合優化和運籌學上。
泰米斯·帕爾帕納斯,法國巴黎大學及法國大學研究所
泰米斯·帕爾帕納斯是法國大學研究所 (IUF) 的高級成員,並擔任法國巴黎大學的計算機科學教授,擔任巴黎數據智能研究所 (diiP) 和數據密集型及知識導向系統 (diNo) 小組的主任。他是兩項法國專利和九項美國專利的作者,其中三項已在全球領先的商業數據管理產品中實施。他曾獲得三項最佳論文獎和 IBM 共享大學研究 (SUR) 獎。他目前在非常大型數據庫 (VLDB) 基金會的董事會任職,擔任 BDR 期刊的主編,並是 IS 期刊的編輯諮詢委員會成員,以及 SIGMOD 2021 的高級程序委員會成員。
目錄大綱
Table of Contents
Entity Resolution: Past, Present, and Yet-to-Come
Preliminaries
Generation 1: Addressing Veracity
Generation 2: Also Addressing Volume
Generation 3: Also Addressing Variety
Generation 4: Also Addressing Velocity
Leveraging External Knowledge
Resources for Entity Resolution
Possible Directions for Future Work
Bibliography
Authors' Biographies
目錄大綱(中文翻譯)
Table of Contents
Entity Resolution: Past, Present, and Yet-to-Come
Preliminaries
Generation 1: Addressing Veracity
Generation 2: Also Addressing Volume
Generation 3: Also Addressing Variety
Generation 4: Also Addressing Velocity
Leveraging External Knowledge
Resources for Entity Resolution
Possible Directions for Future Work
Bibliography
Authors' Biographies