Using CNNs
This project was part of my coursework for the Udacity Machine Learning Engineer nanodegree.
I made use of Amazon's SageMaker, S3 and Lambda services. Meant to perform binary classification (plagiarized or not), the model made use of several computed similarity features to classify texts given a source text to compare against.
The similarity features I used were containment and longest common subsequence between the answer and source texts. Containment was calculated by computing the number of n-grams (where an n-gram is a subset of either text consisting of n words) shared between the answer and source, followed by dividing by the number of n-grams in the answer text. For example, for a answer text which is a word-for-word copy of the source text, but with sentences in a different order, would have a 1-gram containment of 1 (all words in the source appear in the answer, with the answer having the same word count as the source).
The second feature was normalized longest common subsequence, where the length of the longest set of words that appears in both the source and answer (there may be other words in between) is divided by the length of the answer text. This required a dynamic programming approach. Given these 2 features, a random forest predictor had an accuracy of 100% on the test set, but as noted, this was likely due to limited variability in the dataset used.