TY - GEN
T1 - Pirá
AU - Paschoal, André F.A.
AU - Pirozelli, Paulo
AU - Freire, Valdinei
AU - Delgado, Karina V.
AU - Peres, Sarajane M.
AU - José, Marcos M.
AU - Nakasato, Flávio
AU - Oliveira, André S.
AU - Brandão, Anarosa A.F.
AU - Costa, Anna H.R.
AU - Cozman, Fabio G.
N1 - Funding Information:
This work was carried out at the Center for Artificial Intelligence (C4AI-USP), with support by the São Paulo Research Foundation (FAPESP grant #2019/07665-4) and by the IBM Corporation. This research was also partially supported by Itaú Unibanco S.A.; M. M. José, F. Nakasato and A. S. Oliveira have been supported by the Itaú Scholarship Program (PBI) of the Data Science Center (C2D) of the Escola Politécnica da Universidade de São Paulo. A. H. R. Costa and F. G. Cozman thank the support of the National Council for Scientific and Technological Development of Brazil (CNPq grants #310085/2020-9 and #312180/2018-7, respectively).
Publisher Copyright:
© 2021 Owner/Author.
PY - 2021/10/26
Y1 - 2021/10/26
N2 - Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pirá dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pirá is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pirá dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pirá, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.
AB - Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pirá dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pirá is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pirá dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pirá, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.
KW - bilingual dataset
KW - ocean dataset
KW - Portuguese-English dataset
KW - question-answering dataset
UR - http://www.scopus.com/inward/record.url?scp=85119186179&partnerID=8YFLogxK
UR - https://www.mendeley.com/catalogue/439ad74f-06b0-3bfd-a4a3-c4f8da554ff2/
U2 - 10.1145/3459637.3482012
DO - 10.1145/3459637.3482012
M3 - Contribución a la conferencia
AN - SCOPUS:85119186179
SN - 9781450384469
T3 - International Conference on Information and Knowledge Management, Proceedings
SP - 4544
EP - 4553
BT - CIKM 2021 - Proceedings of the 30th ACM International Conference on Information and Knowledge Management
PB - Association for Computing Machinery
Y2 - 1 November 2021 through 5 November 2021
ER -