2024

Unearthing Large Scale Domain-Specific Knowledge from Public Corpora

Fei, Zhaoye, Shao, Yunfan, Li, Linyang et al.

Understand

Large language models (LLMs) have demonstrated remarkable potential in various tasks, however, there remains a significant lack of open-source models and data for specific domains.

  • Previous work has primarily focused on manually specifying resources and collecting high-quality data for specific domains, which is extremely time-consuming and labor-intensive.
  • To address this limitation, we introduce large models into the data collection pipeline to guide the generation of domain-specific information and retrieve relevant data from Common Crawl (CC), a large public corpus.
  • We refer to this approach as Retrieve-from-CC.

Reading the bibliography…