Human or AI? Study Spots Machine-Written and Rephrased Text With 96.46% Accuracy

The researchers treated AI-rephrased writing as a separate category because it carries the ideas that begin with a human, and the vocabulary, sentence structure, or rhythm can reflect the model that rewrites them.

Human or AI? Study Spots Machine-Written and Rephrased Text With 96.46% Accuracy
Representative image Image Credit: ChatGPT

A product review can sound personal without coming from someone who used the product, and a social media post can carry a person's original idea even after AI rewrites every sentence. These blurred boundaries make identifying a text's origins more complicated than checking whether it sounds robotic.

The study "Hybrid deep learning for three-way classification of human-written, AI-generated, and AI-rephrased text," published as an article in press in npj Scientific Reports, explores this problem by asking detection systems to recognise three kinds of writing. Najma Sadia and colleagues report that their best combined system reached 96.46% accuracy on the study's test set, a promising result whose relevance depends on the kinds of text and AI models being tested.

The third category that changes the detection problem

The researchers treated AI-rephrased writing as a separate category because it carries the ideas that begin with a human, and the vocabulary, sentence structure, or rhythm can reflect the model that rewrites them. Detecting that mixture matters for academic integrity, journalism, and the credibility of online reviews, where knowing how content was produced can influence trust.

The dataset contained 161,788 English-language samples, including 46,844 human-written texts, 57,247 AI-generated texts, and 57,697 AI-rephrased texts. Human material came from Amazon product reviews and Twitter posts, giving the experiment two different writing settings.

GPT/ChatGPT, DeepSeek, and Kimi produced the synthetic material. For fresh AI-generated content, the models received instructions without example answers, a method called zero-shot prompting. Rephrasing requests included three examples of original and rewritten sentences, guiding the models to preserve meaning and change expression.

The team divided the corpus into 113,249 training samples, 16,180 validation samples, and 32,359 test samples. The authors describe the dataset and construction protocol as publicly available, giving other researchers a foundation for further testing.

Reading the clues hidden in everyday writing

Researchers measured 68 handcrafted features covering writing style, vocabulary, readability, sentiment, grammar, sentence structure, and authorial patterns. These measurements captured details such as word diversity, clause complexity, and variation in sentence length.

A second set of features came from RoBERTa and the all-MiniLM-L6-v2 Sentence Transformer, which represent relationships between words and their context numerically. The researchers compressed those representations into 100 dimensions and combined them with the handcrafted measurements, creating a 168-feature input for conventional machine learning models.

Random Forest, Support Vector Machine, and Logistic Regression learned from this combined representation. Two transformer models, RoBERTa-base and DeBERTa-v3-base, were trained directly on tokenized text, allowing them to learn contextual patterns without receiving the handcrafted measurements.

Human writing in this corpus showed broader variation in sentence length, vocabulary, and sentiment. Synthetic writing generally occupied a more compact range, with more uniform structures and more concentrated positive sentiment. Those observations describe patterns in this dataset; an upbeat tone or tidy sentence structure cannot establish that an individual passage came from AI.

Adding handcrafted features improved conventional classifiers by roughly 9–10 percentage points compared with using contextual representations alone. Random Forest rose from 70% to about 80%, demonstrating that measurable writing habits contributed useful information beyond the numerical representation of context.

How the strongest system reached 96.46%

The transformer models delivered the strongest individual results: RoBERTa achieved 95.49% accuracy, and DeBERTa-v3 reached 94.88%. The paper's main comparison reports roughly 80% for Random Forest and 74% for both Support Vector Machine and Logistic Regression. One combined approach averaged predictions from the five main models, reaching 95.66% accuracy. Its small improvement over RoBERTa was not clear at the reported 95% confidence level.

The best result came from a method called stacking, in which another classifier learns how to combine predictions from several models. The described stacking system used Random Forest, XGBoost, Support Vector Machine, and Logistic Regression as its base learners, with a Random Forest classifier making the final decision.

That configuration reached 96.46% accuracy, with a reported 95% confidence interval of 96.26% to 96.66%. Alternative final classifiers achieved closely grouped results, ranging from 96.20% to 96.32%. The largest remaining confusion involved distinguishing fresh AI-generated text from AI-rephrased text: the best system labelled 407 AI-generated samples as rephrased and 600 rephrased samples as AI-generated.

Rephrased writing preserves human meaning and acquires machine-shaped expression, giving detectors overlapping signals rather than a clean boundary. The study identifies this category as its most persistent challenge.

Experiments that removed parts of the framework also showed the importance of contextual information and careful training. Handcrafted features alone reached about 62% accuracy, and reducing transformer training from 15 epochs to five lowered accuracy by approximately six percentage points. Removing the learning-rate warmup reduced final validation accuracy by another two to three percentage points.

What these results mean beyond the test set

The detector's 96.46% accuracy applies to this study's test data, which included English reviews and tweets produced using GPT/ChatGPT, DeepSeek, and Kimi. Training and testing used the same three AI model families. Performance on unseen generators, including Claude, Gemini, Llama, and Mistral, remains untested. The experiment left academic essays, news articles, blogs, technical documents, other languages, and mixed-language writing outside its verified scope.

The dataset had unequal category sizes, 106 texts appearing in both training and test sets, and 18 original–AI rewrite pairs spread across different sets. The authors corrected earlier errors in dataset totals. Their estimates of uncertainty used one data split and did not measure how results might change with different training seeds.

The larger systems need substantial computing power, making them better suited to checking documents in batches than providing instant results. Future work includes testing unfamiliar generators and writing domains, expanding into multilingual content, strengthening resistance to paraphrasing, and building smaller models that need less computing power. Detailed explanations of transformer decisions through tools such as SHAP or LIME remain future work.

  • FIRST PUBLISHED IN:
  • Devdiscourse
Give Feedback

Use this form for editorial or site feedback. We usually reply within 2 to 3 working days.

By submitting, you agree that we may use your email address to respond.