Abstract
We release our synthetic parallel paraphrase corpus across 17 languages: Arabic, Catalan, Czech, German, English, Spanish, Estonian, French, Hindi, Indonesian, Italian, Dutch, Romanian, Russian, Swedish, Vietnamese, and Chinese. Our method relies only on monolingual data and a neural machine translation system to generate paraphrases, hence simple to apply. We generate multiple translation samples using beam search and choose the most lexically diverse pair according to their sentence BLEU. We compare our generated corpus with the ParaBank2. According to our evaluation, our synthetic paraphrase pairs are semantically similar and lexically diverse.
| Original language | English |
|---|---|
| Pages | 666-675 |
| Number of pages | 10 |
| Publication status | Published - 2021 |
| Event | 35th Pacific Asia Conference on Language, Information and Computation, PACLIC 2021 - Shanghai, China Duration: 5 Nov 2021 → 7 Nov 2021 |
Conference
| Conference | 35th Pacific Asia Conference on Language, Information and Computation, PACLIC 2021 |
|---|---|
| Country/Territory | China |
| City | Shanghai |
| Period | 5/11/21 → 7/11/21 |
Fingerprint
Dive into the research topics of 'ParaCotta: Synthetic Multilingual Paraphrase Corpora from the Most Diverse Translation Sample Pair'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver