Skip to main navigation Skip to search Skip to main content

ParaCotta: Synthetic Multilingual Paraphrase Corpora from the Most Diverse Translation Sample Pair

  • Alham Fikri Aji*
  • , Tirana Noor Fatyanosa*
  • , Radityo Eko Prasojo*
  • , Philip Arthur
  • , Suci Fitriany*
  • , Salma Qonitah*
  • , Nadhifa Zulfa*
  • , Tomi Santoso*
  • , Mahendra Data
  • *Corresponding author for this work

Research output: Contribution to conferencePaperpeer-review

Abstract

We release our synthetic parallel paraphrase corpus across 17 languages: Arabic, Catalan, Czech, German, English, Spanish, Estonian, French, Hindi, Indonesian, Italian, Dutch, Romanian, Russian, Swedish, Vietnamese, and Chinese. Our method relies only on monolingual data and a neural machine translation system to generate paraphrases, hence simple to apply. We generate multiple translation samples using beam search and choose the most lexically diverse pair according to their sentence BLEU. We compare our generated corpus with the ParaBank2. According to our evaluation, our synthetic paraphrase pairs are semantically similar and lexically diverse.

Original languageEnglish
Pages666-675
Number of pages10
Publication statusPublished - 2021
Event35th Pacific Asia Conference on Language, Information and Computation, PACLIC 2021 - Shanghai, China
Duration: 5 Nov 20217 Nov 2021

Conference

Conference35th Pacific Asia Conference on Language, Information and Computation, PACLIC 2021
Country/TerritoryChina
CityShanghai
Period5/11/217/11/21

Fingerprint

Dive into the research topics of 'ParaCotta: Synthetic Multilingual Paraphrase Corpora from the Most Diverse Translation Sample Pair'. Together they form a unique fingerprint.

Cite this