Resource: Romanian - English news corpus (Processed)

Reference Romanian - English news corpus (Processed)
Date of Submission March 9, 2020, 12:27 p.m.
Status accepted
ISLRN 100-905-126-706-7
Resource Type Primary Text
Media Type Text
Source
Language English, Romanian
Format/MIME Type application/x-tmx+xml
Size 98098 translationUnits, 4908272 words
Access Medium downloadable
Description

This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu.
Bilingual Romanian – English news corpus built from SouthEast European Times (2008 dump). The texts are positionaly aligned, i.e. the sentence on line i in the English text is aligned with the sentence on line i in the Romanian text. Alignment was manually validated.

Version 2.0
Distributor ELRA
Rights Holder Southeast European Times