Mining clinical relationships from patient narratives

Abstract

Background

The Clinical E-Science Framework (CLEF) project has built a system to extract clinically significant information from the textual component of medical records in order to support clinical research, evidence-based healthcare and genotype-meets-phenotype informatics. One part of this system is the identification of relationships between clinically important entities in the text. Typical approaches to relationship extraction in this domain have used full parses, domain-specific grammars, and large knowledge bases encoding domain knowledge. In other areas of biomedical NLP, statistical machine learning (ML) approaches are now routinely applied to relationship extraction. We report on the novel application of these statistical techniques to the extraction of clinical relationships.

Results

We have designed and implemented an ML-based system for relation extraction, using support vector machines, and trained and tested it on a corpus of oncology narratives hand-annotated with clinically important relationships. Over a class of seven relation types, the system achieves an average F1 score of 72%, only slightly behind an indicative measure of human inter annotator agreement on the same task. We investigate the effectiveness of different features for this task, how extraction performance varies between inter- and intra-sentential relationships, and examine the amount of training data needed to learn various relationships.

Conclusion

We have shown that it is possible to extract important clinical relationships from text, using supervised statistical ML techniques, at levels of accuracy approaching those of human annotators. Given the importance of relation extraction as an enabling technology for text mining and given also the ready adaptability of systems based on our supervised learning approach to other clinical relationship extraction tasks, this result has significance for clinical text mining more generally, though further work to confirm our encouraging results should be carried out on a larger sample of narratives and relationship types.

Metadata

Item Type:	Article
Authors/Creators:	Roberts, A. Gaizauskas, R. https://orcid.org/0000-0002-3356-5126 Hepple, M. https://orcid.org/0000-0003-1488-257X Guo, Y.K.
Copyright, Publisher and Additional Information:	© Roberts et al; licensee BioMed Central Ltd. 2008. This article is published under license to BioMed Central Ltd. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Keywords:	information extraction; language system; trees; text
Dates:	Published: 19 November 2008
Institution:	The University of Sheffield
Academic Units:	The University of Sheffield > Faculty of Engineering (Sheffield) > Department of Computer Science (Sheffield)
Depositing User:	Symplectic Sheffield
Date Deposited:	08 Aug 2016 15:18
Last Modified:	23 Jun 2023 22:04
Published Version:	http://dx.doi.org/10.1186/1471-2105-9-S11-S3
Status:	Published
Publisher:	BioMed Central
Refereed:	Yes
Identification Number:	10.1186/1471-2105-9-S11-S3
Open Archives Initiative ID (OAI ID):	oai:eprints.whiterose.ac.uk:99064

CORE (COnnecting REpositories)

Mining clinical relationships from patient narratives

Abstract

Metadata

Download

Published Version

Export

Statistics