Brierley, C, Sawalha, M and Atwell, E (2012) Open-source boundary-annotated corpus for Arabic speech and language processing. In: Chair, NCC, Choukri, K, Declerck, T, an, MUUD, Maegaard, B, Mariani, J, Odijk, J and Piperidis, S, (eds.) Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12). Eighth International Conference on Language Resources and Evaluation (LREC’12), 21-27 May 2012, Istanbul, Turkey. European Language Resources Association (ELRA) , 1011 - 1016. ISBN 978-2-9517408-7-7
Abstract
A boundary-annotated and part-of-speech tagged corpus is a prerequisite for developing phrase break classifiers. Boundary annotations in English speech corpora are descriptive, delimiting intonation units perceived by the listener. We take a novel approach to phrase break prediction for Arabic, deriving our prosodic annotation scheme from Tajwīd (recitation) mark-up in the Qur‟an which we then interpret as additional text-based data for computational analysis. This mark-up is prescriptive, and signifies a widely-used recitation style, and one of seven original styles of transmission. Here we report on version 1.0 of our Boundary-Annotated Qur‟an dataset of 77430 words and 8230 sentences, where each word is tagged with prosodic and syntactic information at two coarse-grained levels. In (Sawalha et al., 2012), we use the dataset in phrase break prediction experiments. This research is part of a larger-scale project to produce annotation schemes, language resources, algorithms, and applications for Classical and Modern Standard Arabic.
Metadata
| Item Type: | Proceedings Paper | 
|---|---|
| Authors/Creators: | 
  | 
        
| Editors: | 
  | 
        
| Keywords: | Prosodic annotation; psycholinguistic chunking; phrase break prediction | 
| Dates: | 
  | 
        
| Institution: | The University of Leeds | 
| Academic Units: | The University of Leeds > Faculty of Engineering & Physical Sciences (Leeds) > School of Computing (Leeds) > Artificial Intelligence & Biological Systems (Leeds) | 
| Depositing User: | Symplectic Publications | 
| Date Deposited: | 02 Dec 2014 15:33 | 
| Last Modified: | 19 Dec 2022 13:29 | 
| Published Version: | http://www.lrec-conf.org/proceedings/lrec2012/pdf/... | 
| Status: | Published | 
| Publisher: | European Language Resources Association (ELRA) | 
| Related URLs: | |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:81620 | 
 CORE (COnnecting REpositories)
 CORE (COnnecting REpositories)