TY - GEN
T1 - Interactive extractive search over biomedical corpora
AU - Taub-Tabib, Hillel
AU - Shlain, Micah
AU - Sadde, Shoval
AU - Lahav, Dan
AU - Eyal, Matan
AU - Cohen, Yaara
AU - Goldberg, Yoav
N1 - Publisher Copyright:
© Association for Computation Linguistics.
PY - 2020/1/1
Y1 - 2020/1/1
N2 - We present a system that allows life-science researchers to search a linguistically annotated corpus of scientific texts using patterns over dependency graphs, as well as using patterns over token sequences and a powerful variant of boolean keyword queries. In contrast to previous attempts to dependency-based search, we introduce a light-weight query language that does not require the user to know the details of the underlying linguistic representations, and instead to query the corpus by providing an example sentence coupled with simple markup. Search is performed at an interactive speed due to efficient linguistic graphindexing and retrieval engine. This allows for rapid exploration, development and refinement of user queries. We demonstrate the system using example workflows over two corpora: the PubMed corpus including 14,446,243 PubMed abstracts and the CORD-19 dataset, a collection of over 45,000 research papers focused on COVID-19 research. The system is publicly available at https://allenai. github.io/spike
AB - We present a system that allows life-science researchers to search a linguistically annotated corpus of scientific texts using patterns over dependency graphs, as well as using patterns over token sequences and a powerful variant of boolean keyword queries. In contrast to previous attempts to dependency-based search, we introduce a light-weight query language that does not require the user to know the details of the underlying linguistic representations, and instead to query the corpus by providing an example sentence coupled with simple markup. Search is performed at an interactive speed due to efficient linguistic graphindexing and retrieval engine. This allows for rapid exploration, development and refinement of user queries. We demonstrate the system using example workflows over two corpora: the PubMed corpus including 14,446,243 PubMed abstracts and the CORD-19 dataset, a collection of over 45,000 research papers focused on COVID-19 research. The system is publicly available at https://allenai. github.io/spike
UR - https://www.scopus.com/pages/publications/85095302249
M3 - Conference contribution
AN - SCOPUS:85095302249
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 28
EP - 37
BT - BioNLP 2020 - 19th SIGBioMed Workshop on Biomedical Language Processing, Proceedings of the Workshop
PB - Association for Computational Linguistics (ACL)
T2 - 19th SIGBioMed Workshop on Biomedical Language Processing, BioNLP 2020 at the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020
Y2 - 9 July 2020
ER -