From centralized to ad-hoc knowledge base construction for hypotheses generation

Shaked Launer-Wachs, Hillel Taub-Tabib, Jennie Tokarev Madem, Orr Bar-Natan, Yoav Goldberg, Yosi Shamay

Research output: Contribution to journalArticlepeer-review

2 Scopus citations

Abstract

Objective: To demonstrate and develop an approach enabling individual researchers or small teams to create their own ad-hoc, lightweight knowledge bases tailored for specialized scientific interests, using text-mining over scientific literature, and demonstrate the effectiveness of these knowledge bases in hypothesis generation and literature-based discovery (LBD). Methods: We propose a lightweight process using an extractive search framework to create ad-hoc knowledge bases, which require minimal training and no background in bio-curation or computer science. These knowledge bases are particularly effective for LBD and hypothesis generation using Swanson's ABC method. The personalized nature of the knowledge bases allows for a somewhat higher level of noise than “public facing” ones, as researchers are expected to have prior domain experience to separate signal from noise. Fact verification is shifted from exhaustive verification of the knowledge base to post-hoc verification of specific entries of interest, allowing researchers to assess the correctness of relevant knowledge base entries by considering the paragraphs in which the facts were introduced. Results: We demonstrate the methodology by constructing several knowledge bases of different kinds: three knowledge bases that support lab-internal hypothesis generation: Drug Delivery to Ovarian Tumors (DDOT); Tissue Engineering and Regeneration; Challenges in Cancer Research; and an additional comprehensive, accurate knowledge base designated as a public resource for the wider community on the topic of Cell Specific Drug Delivery (CSDD). In each case, we show the design and construction process, along with relevant visualizations for data exploration, and hypothesis generation. For CSDD and DDOT we also show meta-analysis, human evaluation, and in vitro experimental evaluation. Conclusion: Our approach enables researchers to create personalized, lightweight knowledge bases for specialized scientific interests, effectively facilitating hypothesis generation and literature-based discovery (LBD). By shifting fact verification efforts to post-hoc verification of specific entries, researchers can focus on exploring and generating hypotheses based on their expertise. The constructed knowledge bases demonstrate the versatility and adaptability of our approach to versatile research interests. The web-based platform, available at https://spike-kbc.apps.allenai.org, provides researchers with a valuable tool for rapid construction of knowledge bases tailored to their needs.

Original languageEnglish
Article number104383
JournalJournal of Biomedical Informatics
Volume142
DOIs
StatePublished - Jun 2023

Bibliographical note

Publisher Copyright:
© 2023 Elsevier Inc.

Funding

This work was supported by the Israel Science Foundation [ISF grant #901/91 ] and ERC Starting Grant #802774 (iEXTRACT).

FundersFunder number
European Research Council802774
Israel Science Foundation901/91

    Keywords

    • Extractive search
    • Hypothesis generation
    • Knowledge base
    • Literature-based discovery
    • Rapid exploration

    Fingerprint

    Dive into the research topics of 'From centralized to ad-hoc knowledge base construction for hypotheses generation'. Together they form a unique fingerprint.

    Cite this