Caio Alencar-Palha, Thaís Ocampo, Thaísa Pinheiro Silva, Deborah Queiroz Freitas, Frederico Sampaio Neves, Matheus Lima Oliveira
5 min
As generative AI becomes more prevalent in academic writing, researchers are increasingly using "humanization" tools to bypass AI content detectors. This study investigates whether these dedicated humanization platforms and chatbot-based rephrasing techniques can successfully mask AI-generated scientific abstracts in the field of oral radiology.
The researchers selected 30 high-impact scientific articles from dentomaxillofacial radiology journals. They used the Methodology and Results sections of these articles to prompt two chatbots (ChatGPT and Gemini) to generate complete abstracts. These AI-generated abstracts were then processed through four different "humanization" methods: rephrasing via ChatGPT, rephrasing via Gemini, and using two dedicated platforms, HumanizeAI and StealthWriter. The team then used GPTZero to analyze all 330 resulting abstracts to determine the percentage of AI-generated content, comparing these scores against original human-written abstracts.
The study found that while raw AI-generated abstracts are easily identified by GPTZero, the use of dedicated humanization platforms significantly lowers detection rates to levels statistically indistinguishable from human-written text. In contrast, asking chatbots to rewrite their own content failed to reduce detection rates and, in some cases, actually increased them. The results highlight a growing "arms race" between AI generation tools and detection software, suggesting that reliance on automated detectors for academic integrity is increasingly precarious.
These findings demonstrate that current AI detection tools are vulnerable to specialized rephrasing platforms, which could allow for the undetected submission of AI-generated work. The authors argue that educators and journal editors cannot rely solely on software to police academic integrity. Instead, they must shift toward more robust assessment methods—such as oral examinations and critical reflection—and establish clear, transparent policies regarding the ethical use of AI in scientific publishing.
The rise of generative artificial intelligence (AI) in scientific writing has prompted the development of platforms designed to "humanize" AI-generated text by rephrasing content to reduce its detectability as machine-generated. This study aimed to evaluate the performance of chatbots and dedicated platforms in humanizing AI-generated abstracts. Thirty high-impact scientific articles published in dentomaxillofacial radiology journals were selected. The Methodology and Results sections of each abstract were used to prompt two chatbots—ChatGPT and Gemini—to generate complete abstracts. These AI-generated abstracts were then humanized using the same chatbots along with two dedicated platforms: HumanizeAI and StealthWriter. The percentage of AI-generated content was determined for all abstracts (n = 330), including the original human-written abstracts, using GPTZero. AI detection rates were compared using two-way analysis of variance followed by Tukey's post hoc test (α = 0.05). The results showed that original abstracts and AI-generated abstracts humanized by HumanizeAI and StealthWriter exhibited significantly lower GPTZero detection rates (p < 0.05), whereas abstracts generated or humanized by chatbots showed significantly higher detection rates (p < 0.05). In conclusion, chatbot-generated scientific abstracts humanized by dedicated platforms exhibit low AI detection rates, whereas those humanized by chatbots exhibit considerably higher detection levels. Educators must prepare students with critical appraisal skills to identify AI-generated content and uphold academic integrity, ensuring that AI serves as a complementary tool rather than a means of misrepresentation.
Alex: Which is the logical endpoint of this approach—the humanizer has learned to exploit the detector's own decision boundary and push output to the far side of it. It's not just evading detection; it's inverting the false positive rate.
Sam: That's a meaningful result. But how much does it depend on the specific detector they used?
Alex: That's the central limitation, and a careful referee would push back hard here. The study relied solely on GPTZero. Whether this generalizes to other detectors—or to longer, structurally complex manuscripts where coherence is harder to fake than in a 250-word abstract—is an open question the paper doesn't resolve.
Sam: So the "arms race is over" framing might be overstated if you're thinking about full-length papers rather than abstracts?
Alex: It's a fair caveat. The abstract is the easiest unit to manipulate—short, self-contained, stylistically constrained. A full methods section or a results narrative with internal logical dependencies is a harder target. The paper doesn't test that, and the authors don't claim it.
Sam: There's also a finding about prompt engineering versus dedicated platforms, right? Because the obvious workaround is just asking the LLM to sound more human.
Alex: And that's where the results are instructive. When ChatGPT and Gemini were simply prompted to rewrite their own output in a more natural, human style, detection rates didn't fall nearly as far. The dedicated platforms outperformed that approach by a wide margin—which tells you they're doing something architecturally different. They're not applying a stylistic overlay; they're specifically targeting the statistical features that detectors use as their signal.
Sam: So the gap between "ask the chatbot to humanize it" and "use a purpose-built tool" is real and measurable.
Alex: It is, and that distinction matters for how we think about the threat model. This isn't about sophisticated prompt engineering that any user could replicate. It's about tools specifically optimized to defeat a particular class of classifier.
Sam: If content-based detection is becoming unreliable, where does that leave editorial practice?
Alex: The authors' position is that the field needs to shift from analyzing text to verifying provenance. Cryptographic watermarking, version-history tracking, submission metadata—approaches that establish how a document was produced rather than trying to infer it from style. The argument is that trying to catch AI by reading its statistical fingerprint is a game the technology has already learned to win.
Sam: It's a shift from forensic analysis to chain-of-custody verification.
Alex: That's a good way to frame it. And it has real implications for how journals structure their submission workflows, not just which detection software they license. Thanks for listening to ResearchPod.