Zhang, H., Chen, Q., Zou, Y. et al. (3 more authors) (2024) Document set expansion with positive-unlabelled learning using intractable density estimation. In: Calzolari, N., Kan, M-Y, Hoste, V., Lenci, A., Sakti, S. and Xue, N., (eds.) Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). The 2024 joint international conference on computational linguistics, language resources and evaluation, 20-25 May 2024, Torino, Italia. ELRA and ICCL , pp. 5167-5173.
Abstract
The Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unrealistic assumption of knowing the class prior for positive samples in the collection. To address this limitation, this paper introduces a novel PU learning framework that utilizes intractable density estimation models. Experiments conducted on PubMed and Covid datasets in a transductive setting showcase the effectiveness of the proposed method for DSE. Code is available from https://github.com/Beautifuldog01/Document-set-expansion-puDE.
Metadata
Item Type: | Proceedings Paper |
---|---|
Authors/Creators: |
|
Editors: |
|
Copyright, Publisher and Additional Information: | © 2024 The author(s). Except as otherwise noted, this author-accepted version of a paper published in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) is made available via the University of Sheffield Research Publications and Copyright Policy under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution and reproduction in any medium, provided the original work is properly cited. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/ © 2024 ELRA Language Resource Association: This is an Open Access paper distributed under the terms of the Creative Commons Attribution-NonCommercial Licence (https://creativecommons.org/licenses/by-nc/4.0/). |
Keywords: | Document set expansion; PU learning, Information retrieval; Density estimation |
Dates: |
|
Institution: | The University of Sheffield |
Academic Units: | The University of Sheffield > Faculty of Engineering (Sheffield) > Department of Computer Science (Sheffield) |
Depositing User: | Symplectic Sheffield |
Date Deposited: | 16 Apr 2024 14:35 |
Last Modified: | 21 May 2024 11:54 |
Published Version: | https://aclanthology.org/2024.lrec-main.460 |
Status: | Published |
Publisher: | ELRA and ICCL |
Refereed: | Yes |
Related URLs: | |
Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:211557 |
Downloads
Filename: 732_Paper (1).pdf
Licence: CC-BY 4.0
Filename: 2024.lrec-main.460.pdf
Licence: CC-BY-NC 4.0