ResearchPod Summary
Automated Essay Scoring (AES) systems typically require large amounts of human-scored data, which is expensive to collect. While synthetic data can mitigate this, conventional generation often produces overly polished essays that fail to capture the grammatical errors and proficiency-related patterns characteristic of language learners. This paper investigates whether explicitly supervising an LLM generator with error-annotated learner writing—typically used for Grammatical Error Detection (GED)—can produce more realistic synthetic essays that improve downstream AES performance.
The author fine-tuned a Qwen2-7B-Instruct model using Quantized Low-Rank Adaptation (QLoRA) on three learner corpora: CLC FCE, Write & Improve, and ELLIPSE. The model was trained to generate essays with inline error tags (e.g., identifying missing determiners or tense errors). These tags were then removed to create erroneous surface essays for training BERT-based AES scorers. The study compares this proposed approach against two baselines: authentic human-written essays and conventionally generated synthetic essays (without error supervision).
In larger-data settings, the proposed error-supervised synthetic data consistently outperformed the conventional synthetic baseline across nearly all evaluation metrics. On the CLC FCE dataset, the proposed approach even achieved performance comparable to models trained on authentic human data. Qualitative and quantitative analyses confirmed that the generated essays successfully reproduced learner-like error distributions, such as tense inconsistencies and missing articles, which are often absent in standard synthetic generation.
However, the results in extremely low-resource settings (50-100 training samples) were mixed and unstable. While the proposed approach showed advantages at 200 training samples, it did not consistently outperform the conventional baseline at the smallest sample sizes. Furthermore, while the approach is effective, the generated essays still suffer from occasional repetition, semantic inconsistencies, and over-correction, indicating that they improve upon, but do not perfectly replicate, authentic learner writing.
This study demonstrates that improving the quality of synthetic learner data does not necessarily require complex post-hoc error-injection pipelines. Instead, a simple change to the training objective—incorporating existing error annotations—can yield more effective training data for AES. Importantly, the experiments on the ELLIPSE corpus suggest that this method is not limited to datasets with manual error annotations; automatically derived error tags (using GEC/GED tools) can also serve as a viable supervision signal, significantly broadening the applicability of this approach to a wider range of educational datasets.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.