A Multi-Similarity
Neural Network for Paraphrase Detection
This study introduces a multi-similarity neural
network framework for paraphrase detection, an important task in natural
language processing that identifies whether two sentences convey the same
meaning using different expressions. The proposed method combines various
similarity measures, such as string-based similarity, semantic similarity, and
embedding-based similarity, with a deep learning classifier. The framework is
structured as a three-phase pipeline: preprocessing, extraction of multiple
similarity features, and classification through a neural network. It employs
more than 168 string similarity algorithms, semantic measures derived from
WordNet, and several pre-trained embedding models to compute similarity scores.
These features are aggregated and supplied to a deep neural network to
determine whether sentence pairs are paraphrases. The model was evaluated on
the Microsoft Research Paraphrase Corpus (MSRP) using accuracy and F1-score as
performance metrics. The experimental results indicate that the proposed
framework achieves 81.74% accuracy and an F1 Score of 86.6%, surpassing several
existing approaches. Overall, the results suggest that integrating diverse
similarity measures with neural networks enhances the identification of both
explicit and nuanced paraphrases, thereby supporting advancements in text
analysis and plagiarism detection systems.
[1] Mihalcea, R.; Corley, C.; Strapparava, C. Corpus-based and knowledge-based measures of text semantic similarity. In Proceedings of the AAAI, July 2006, Vol. 6, pp. 775–780.
[2] Hassan,
S. Measuring semantic relatedness using salient encyclopedic concepts. PhD thesis, University of North Texas, 2011.
[3] Rus,
V.; McCarthy, P.M.; Lintean, M.C.; McNamara, D.S.; Graesser, A.C. Paraphrase
Identification with Lexico-Syntactic Graph Subsumption. In Proceedings of the FLAIRS conference, May 2008, pp. 201–206.
[4] Islam,
A.; Inkpen, D. Semantic similarity of short texts. Recent Advances in Natural Language
Processing V 2009, 309, 227–236.
[5] Milajevs,
D.; Kartsaklis, D.; Sadrzadeh, M.; Purver, M. Evaluating neural word
representations in tensor-based compositional settings. arXiv preprint
arXiv:1408.6179, 2014.
[6] Fernando,
S.; Stevenson, M. A semantic similarity
approach to paraphrase detection. In
Proceedings of the Proceedings of the 11th annual research colloquium of the UK
special interest group for computational linguistics, March 2008, pp. 45–52.
[7] Kozareva,
Z.; Montoyo, A. Paraphrase identification on the basis of supervised machine
learning techniques. In Proceedings of the International conference on natural
language processing (in Finland). Springer, August 2006, pp. 524–533.
[8] Qiu,
L.; Kan, M.Y.; Chua, T.S. Paraphrase recognition via dissimilarity significance
classification. In Proceedings of the Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing, July 2006, pp. 18–26.
[9] Ul-Qayyum,
Z.; Altaf, W. Paraphrase identification using semantic heuristic features.
Research Journal of Applied Sciences, Engineering and Technology 2012, 4, 4894–4904.
[10] Blacoe,
W.; Lapata, M. A comparison of vector-based representations for semantic
composition. In Proceedings of the Proceedings of the 2012 joint conference on empirical methods in natural language processing and computational natural
language learning, July 2012, pp. 546–556.
[11] Ji,
Y.; Eisenstein, J. Discriminative improvements to distributional sentence
similarity. In 367 Proceedings of the Proceedings of the 2013 conference on
empirical methods in natural language 368 processing, October 2013, pp.
891–896.
[12] El
Desouki, M.I.; Gomaa, W.H.; Abdalhakim, H. A hybrid model for paraphrase
detection combines pros of text similarity with deep learning. Int. J. Comput.
Appl 2019, 975, 8887.
[13] Finch,
A.; Hwang, Y.S.; Sumita, E. Using machine translation evaluation techniques to
determine sentence-level semantic equivalence. In Proceedings of the Proceedings of the third international workshop on paraphrasing (IWP2005),
2005.
[14] Wan,
S.; Dras, M.; Dale, R.; Paris, C. Using dependency-based features to take
the’para-farce’out of paraphrase. In
Proceedings of the Proceedings of the Australasian language technology workshop
2006, November 2006, pp. 131–138.
[15] Socher, R.; Huang, E.H.; Pennington, J.; Ng, A.Y.; Manning, C.D. Dynamic pooling and unfolding recursive autoencoders for paraphrase detection. In Proceedings of the NIPS, December 2011, Vol. 24, pp. 801–809.
[16] Madnani,
N.; Tetreault, J.; Chodorow, M. Re-examining machine translation metrics for
paraphrase identification. In Proceedings of the Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational
Linguistics: Human Language Technologies, June 2012, pp. 182–190.
[17] He,
H.; Gimpel, K.; Lin, J. Multi-perspective sentence similarity modeling with
convolutional neural networks. In Proceedings of the Proceedings of the 2015 conference on empirical methods in natural language processing, September 2015,
pp. 1576–1586.
[18] Filice,
S.; Da San Martino, G.; Moschitti, A. Structural representations for learning
relations between pairs of texts. In Proceedings of the Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th
International Joint Conference on Natural Language Processing (Volume 1: Long Papers), July 2015, pp. 1003–1013.
[19]
Cheng, J.; Kartsaklis, D. Syntax-aware multi-sense word embeddings for deep
compositional models of meaning. arXiv preprint arXiv:1508.02354 2015.
[20]
Shahmohammadi, H.; Dezfoulian, M.H.; Mansoorizadeh, M. Paraphrase detection
using LSTM networks and handcrafted features. Multimedia Tools and Applications
2021, 80, 6479–6492.
[21] Kong,
L.; Han, Z.; Han, Y.; Qi, H. A deep paraphrase identification model interacting
semantics with syntax. Complexity 2020.
[22] Hany,
M.; Gomaa, W.H. A Hybrid Approach to Paraphrase Detection Based on Text
Similarities and Machine Learning Classifiers. In Proceedings of the 2022 2nd International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC).
IEEE, May 2022, pp. 343–348.
[23] Reimers,
N.; Gurevych, I. Sentence-bert: Sentence embeddings using siamese
bert-networks. arXiv preprint arXiv:1908.10084 2019.
[24] Reimers,
N.; Gurevych, I. Making monolingual
sentence embeddings multilingual using knowledge distillation. arXiv preprint
arXiv:2004.09813 2020.
[25] Feng,
F.; Yang, Y.; Cer, D.; Arivazhagan, N.; Wang, W. Language-agnostic BERT sentence embedding.
arXiv preprint arXiv:2007.01852 2020.
[26] Bojanowski,
P.; Grave, E.; Joulin, A.; Mikolov, T.
Enriching word vectors with subword information. Transactions of the
association for computational linguistics 2017, 5, 135–146.
[27] Joulin,
A.; Grave, E.; Bojanowski, P.; Mikolov, T. Bag of tricks for efficient text
classification. arXiv preprint arXiv:1607.01759 2016.
[28] Joulin,
A.; Grave, E.; Bojanowski, P.; Douze, M.; Jégou, H.; Mikolov, T. Fasttext. zip:
Compressing text classification models. arXiv preprint arXiv:1612.03651 2016.
[29] El
Desouki, M.I.; Gomaa, W.H. Exploring the recent trends of paraphrase detection.
International Journal of Computer Applications 2019, 975.