Islamic University Journal of Applied Sciences

Assessing Transformer Models for Abstractive Summarization of Scientific Articles

Emad Nabil

Keywords: Abstractive summarization; Text summarization; Pre-trained Models; CL-SciSumm 2019

Major: Engineering

Sub Major: Computer Science

https://doi.org/10.63070/jesc.2026.033 Received 06 April 2026; Revised 20 May 2026; Accepted 10 June 2026; Available online 15 July 2026.
DownloadPDF
Abstract

The rapid growth of academic literature has intensified the need for effective automatic text summarization techniques capable of producing concise and informative representations of scientific documents. While extractive methods are widely used, they are limited in their ability to generate coherent and semantically rich summaries. Recent advances in Transformer-based architectures have enabled significant progress in abstractive summarization; however, their effectiveness on domain-specific datasets, such as scientific articles, remains an open challenge. In this study, we investigate the performance of three pre-trained Transformer-based models—T5, BART, and GPT-2—on the task of abstractive summarization using the CL-SciSumm 2019 dataset. A total of 19 experimental configurations are conducted to analyze the impact of generation parameters, including beam size, length penalties, and n-gram constraints, on summarization quality. The models are evaluated using ROUGE metrics, with a focus on ROUGE-2.To complement content-based evaluation, this work incorporates linguistic acceptability assessment using the Corpus of Linguistic Acceptability (CoLA), a benchmark dataset for evaluating grammatical correctness. The results show that BART achieves the best performance with an ROUGE-2 F1-score of 0.40664, while T5 demonstrates superior grammatical acceptability, achieving 93.36%, but BART achieves a very near performance to T5. Ultimately, these findings demonstrate the potential of pre-trained neural networks, particularly the BART architecture, to drive the future of complex, generative NLP applications, transforming how academic research is processed and understood.

References

[1] R. R. Waly and W. H. Gomaa, "Extractive Summarization of Scientific Articles," in 2022 2nd International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), May 2022, pp. 349-354. doi: 10.1109/MIUCC55081.2022.9781702.

[2] W. S. El-Kassas, C. R. Salama, A. A. Rafea, and H. K. Mohamed, "Automatic text summarization: A comprehensive survey," Expert Systems with Applications, vol. 165, p. 113679, 2021. doi: 10.1016/j.eswa.2020.113679.

[3] Y. M. Wazery, M. E. Saleh, A. Alharbi, and A. A. Ali, "Abstractive Arabic Text Summarization Based on Deep Learning," Computational Intelligence and Neuroscience, vol. 2022, 2022. doi: 10.1155/2022/9046638.

[4] N. I. Altmami and M. E. B. Menai, "Automatic summarization of scientific articles: A survey," Journal of King Saud University-Computer and Information Sciences, vol. 34, no. 4, pp. 1011-1028, 2020. doi: 10.1016/j.jksuci.2020.04.020.

[5] P. Semaan, "Natural language generation: an overview," J Comput Sci Res, vol. 1, no. 3, pp. 50-57, 2012. doi: 10.14355/jcsr.2012.0103.02.

[6] S. Teufel and M. Moens, "Summarizing scientific articles: experiments with relevance and rhetorical status," Computational Linguistics, vol. 28, no. 4, pp. 409-445, 2002. doi: 10.1162/089120102762671936.

[7] N. Saini, S. Kumar, S. Saha, and P. Bhattacharyya, "Scientific Document Summarization using Citation Context and Multi-objective Optimization," in 2020 25th International Conference on Pattern Recognition (ICPR), Jan. 2021, pp. 4290-4295. doi: 10.1109/ICPR48806.2021.9412089.

[8] C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, "Exploring the limits of transfer learning with a unified text-to-text transformer," J. Mach. Learn. Res., vol. 21, no. 1, Art. no. 140, Jan. 2020. doi: 10.5555/3455716.3455856.

[9] C. Zerva, M. Q. Nghiem, N. T. Nguyen, and S. Ananiadou, "Nactem-uom@ cl-scisumm 2019," in Proceedings of the 4th Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL@ SIGIR), Jan. 2019.

[10] A. Cohan and N. Goharian, "Scientific document summarization via citation contextualization and scientific discourse," International Journal on Digital Libraries, vol. 19, no. 2, pp. 287-303, 2018. doi: 10.1007/s00799-017-0216-7.

[11] M. La Quatra, L. Cagliero, and E. Baralis, "Poli2sum@ cl-scisumm-19: Identify, classify, and summarize cited text spans by means of ensembles of supervised models," in Proceedings of the 4th Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL@ SIGIR), Jan. 2019.

[12] S. Ma, H. Zhang, T. Xu, J. Xu, S. Hu, and C. Zhang, "IR&TM-NJUST@ CLSciSumm-19," in Proceedings of the 4th Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL@ SIGIR), 2019, pp. 181-195.

[13] L. Li, Y. Zhu, Y. Xie, Z. Huang, W. Liu, X. Li, and Y. Liu, "CIST@ CLSciSumm-19: Automatic Scientific Paper Summarization with Citances and Facets," in Proceedings of the 4th Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL@ SIGIR), vol. 54, 2019.

[14] L. Chiruzzo, A. Bravo, and H. Saggion, "LaSTUS-TALN+ INCO@ CL-SciSumm 2019," in Proceedings of the 4th Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL@ SIGIR), Jan. 2019

[15] C. Zerva, M. Q. Nghiem, N. T. Nguyen, and S. Ananiadou, "Cited text span identification for scientific summarisation using pre-trained encoders," Scientometrics, vol. 125, no. 3, pp. 3109-3137, 2020. doi: 10.1007/s11192-020-03741-z.

[16] A. R. Fabbri, W. Kry?ci?ski, B. McCann, C. Xiong, R. Socher, and D. Radev, "Summeval: Re-evaluating summarization evaluation," Transactions of the Association for Computational Linguistics, vol. 9, pp. 391-409, 2021. doi: 10.1162/tacl_a_00373.

[17] M. Lewis et al., "Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension," arXiv preprint arXiv:1910.13461, 2019. doi: 10.48550/arXiv.1910.13461.

[18] A. Radford et al., "Language models are unsupervised multitask learners," OpenAI blog, 2019.

[19] A. Warstadt, A. Singh, and S. R. Bowman, "Neural network acceptability judgments," Transactions of the Association for Computational Linguistics, vol. 7, pp. 625-641, 2019. doi: 10.1162/tacl_a_00290.

[20] R. Haruna, A. Obiniyi, M. Abdulkarim and A. A. Afolorunsho, "Automatic Summarization of Scientific Documents Using Transformer Architectures: A Review," 2022 5th Information Technology for Education and Development (ITED), Abuja, Nigeria, 2022, pp. 1-6, doi: 10.1109/ITED56637.2022.10051602.