Skip to main navigation Skip to search Skip to main content

Paraphrase type identification for plagiarism detection using contexts and word embeddings

  • Faisal Alvi*
  • , Mark Stevenson
  • , Paul Clough
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

30 Scopus citations

Abstract

Paraphrase types have been proposed by researchers as the paraphrasing mechanisms underlying acts of plagiarism. Synonymous substitution, word reordering and insertion/deletion have been identified as some of the common paraphrasing strategies used by plagiarists. However, similarity reports generated by most plagiarism detection systems provide a similarity score and produce matching sections of text with their possible sources. In this research we propose methods to identify two important paraphrase types – synonymous substitution and word reordering in paraphrased, plagiarised sentence pairs. We propose a three staged approach that uses context matching and pretrained word embeddings for identifying synonymous substitution and word reordering. Our proposed approach indicates that the use of Smith Waterman Algorithm for Plagiarism Detection and ConceptNet Numberbatch pretrained word embeddings produces the best performance in terms of F 1 scores. This research can be used to complement similarity reports generated by currently available plagiarism detection systems by incorporating methods to identify paraphrase types for plagiarism detection.

Original languageEnglish
Article number42
JournalRUSC Universities and Knowledge Society Journal
Volume18
Issue number1
DOIs
StatePublished - Dec 2021

Bibliographical note

Publisher Copyright:
© 2021, The Author(s).

Keywords

  • Context matching
  • Paraphrase types
  • Plagiarism
  • Plagiarism detection
  • Synonymous substitution
  • Word embeddings
  • Word reordering

ASJC Scopus subject areas

  • Education
  • Computer Science Applications

Fingerprint

Dive into the research topics of 'Paraphrase type identification for plagiarism detection using contexts and word embeddings'. Together they form a unique fingerprint.

Cite this