Table of Contents
Fetching ...

Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review

Zhyar Rzgar K. Rostam, Gábor Kertész

TL;DR

This systematic literature review analyzes 41 studies (2018–January 2024) on pre-trained language models for domain-specific text classification, using PRISMA-guided methods. It provides a taxonomy of domain-specific PLMs (e.g., BioBERT, SciBERT, MatSciBERT, RadBERT, HumBERT, SsciBERT, NukeBERT) and techniques (transfer learning, prompt-based learning, activation fine-tuning), plus a biomedical case study comparing BioBERT, SciBERT, and BERT. The review finds that domain-adapted PLMs consistently achieve state-of-the-art results across biomedical, financial, nuclear, humanitarian, social science, and materials science domains, while highlighting challenges such as data scarcity, high computational costs, and domain adaptation. It also surveys contemporary techniques (prompting, CARP, in-context learning) and conducts a cross-domain comparison to illustrate performance variability. The study offers guidance for practitioners on model selection and future directions toward efficient, ethical, and data-efficient domain-specific NLP.

Abstract

The exponential increase in scientific literature and online information necessitates efficient methods for extracting knowledge from textual data. Natural language processing (NLP) plays a crucial role in addressing this challenge, particularly in text classification tasks. While large language models (LLMs) have achieved remarkable success in NLP, their accuracy can suffer in domain-specific contexts due to specialized vocabulary, unique grammatical structures, and imbalanced data distributions. In this systematic literature review (SLR), we investigate the utilization of pre-trained language models (PLMs) for domain-specific text classification. We systematically review 41 articles published between 2018 and January 2024, adhering to the PRISMA statement (preferred reporting items for systematic reviews and meta-analyses). This review methodology involved rigorous inclusion criteria and a multi-step selection process employing AI-powered tools. We delve into the evolution of text classification techniques and differentiate between traditional and modern approaches. We emphasize transformer-based models and explore the challenges and considerations associated with using LLMs for domain-specific text classification. Furthermore, we categorize existing research based on various PLMs and propose a taxonomy of techniques used in the field. To validate our findings, we conducted a comparative experiment involving BERT, SciBERT, and BioBERT in biomedical sentence classification. Finally, we present a comparative study on the performance of LLMs in text classification tasks across different domains. In addition, we examine recent advancements in PLMs for domain-specific text classification and offer insights into future directions and limitations in this rapidly evolving domain.

Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review

TL;DR

This systematic literature review analyzes 41 studies (2018–January 2024) on pre-trained language models for domain-specific text classification, using PRISMA-guided methods. It provides a taxonomy of domain-specific PLMs (e.g., BioBERT, SciBERT, MatSciBERT, RadBERT, HumBERT, SsciBERT, NukeBERT) and techniques (transfer learning, prompt-based learning, activation fine-tuning), plus a biomedical case study comparing BioBERT, SciBERT, and BERT. The review finds that domain-adapted PLMs consistently achieve state-of-the-art results across biomedical, financial, nuclear, humanitarian, social science, and materials science domains, while highlighting challenges such as data scarcity, high computational costs, and domain adaptation. It also surveys contemporary techniques (prompting, CARP, in-context learning) and conducts a cross-domain comparison to illustrate performance variability. The study offers guidance for practitioners on model selection and future directions toward efficient, ethical, and data-efficient domain-specific NLP.

Abstract

The exponential increase in scientific literature and online information necessitates efficient methods for extracting knowledge from textual data. Natural language processing (NLP) plays a crucial role in addressing this challenge, particularly in text classification tasks. While large language models (LLMs) have achieved remarkable success in NLP, their accuracy can suffer in domain-specific contexts due to specialized vocabulary, unique grammatical structures, and imbalanced data distributions. In this systematic literature review (SLR), we investigate the utilization of pre-trained language models (PLMs) for domain-specific text classification. We systematically review 41 articles published between 2018 and January 2024, adhering to the PRISMA statement (preferred reporting items for systematic reviews and meta-analyses). This review methodology involved rigorous inclusion criteria and a multi-step selection process employing AI-powered tools. We delve into the evolution of text classification techniques and differentiate between traditional and modern approaches. We emphasize transformer-based models and explore the challenges and considerations associated with using LLMs for domain-specific text classification. Furthermore, we categorize existing research based on various PLMs and propose a taxonomy of techniques used in the field. To validate our findings, we conducted a comparative experiment involving BERT, SciBERT, and BioBERT in biomedical sentence classification. Finally, we present a comparative study on the performance of LLMs in text classification tasks across different domains. In addition, we examine recent advancements in PLMs for domain-specific text classification and offer insights into future directions and limitations in this rapidly evolving domain.
Paper Structure (54 sections, 10 figures, 13 tables)

This paper contains 54 sections, 10 figures, 13 tables.

Figures (10)

  • Figure 1: An overview of this SLR
  • Figure 2: The PRISMA 2020 flow diagram of the performed SLR.
  • Figure 3: Text classification basic steps.
  • Figure 4: Pre-train and Fine-tune with LLM.
  • Figure 5: Comparison of biomedical PLMs on NER tasks. Models on left (blue) are evaluated using standard F1 scores on BC5CDR, BC5-chem, and ShARe13. Models on the right (orange) are evaluated using micro-averaged F1 on the BLURB benchmark.
  • ...and 5 more figures