Provenance of HAVING Queries in Semirings with Monus by Prof Pierre Senellart
Abstract:
The semiring framework and its extensions form the basis of a rich collection of theoretical results and implementations for provenance tracking of database queries. Many real-world queries use aggregation and conditions on the aggregate values. Support for such queries has been proposed by introducing semimodule elements as aggregate values and formal comparisons between aggregate values as tuple annotations, which takes the approach outside the standard semiring framework. In this work, we show how to introduce a semantics for the provenance of such queries in arbitrary commutative semirings with monus (or m-semirings), without the need for additional operators. This semantics is shown to be compatible with the natural join-based rewriting of HAVING COUNT(*) queries in semirings that are absorptive and where times distributes over monus. We derive algorithms for this semantics and show that they can be implemented within the ProvSQL provenance-tracking system with viable performance on a real-world dataset for probabilistic query evaluation.
Biography:
Pierre Senellart is a Professor in the Computer Science Department at the École normale supérieure (ENS-PSL) in Paris, France, as well as Vice-President of PSL University in charge of Digital Infrastructure and IT Convergence. He is an alumnus of ENS-PSL and obtained his M.Sc. (2003) and Ph.D. (2007) in Computer Science from Université Paris-Sud, studying under the supervision of Serge Abiteboul. Before joining ENS, he was an Associate Professor (2008–2013) and then a Professor (2014–2016) at Télécom Paris. He also held secondary appointments as Lecturer at the University of Hong Kong in 2012–2013 and as Senior Research Fellow at the National University of Singapore from 2014 to 2016. Pierre Senellart's research interests focus on the practical and theoretical aspects of Web data management, including Web crawling and archiving, Web information extraction, uncertainty and provenance management, Web mining, and intensional data management.