AI vs. traditional grading: Small LLMs show promise in argument assessment

Writing structured and coherent essays is a fundamental skill for students, but many struggle to develop strong argumentation due to limited individualized feedback. Traditional assessment methods rely on educators manually reviewing essays, which is time-consuming and constrained by high student-to-teacher ratios.

AI vs. traditional grading: Small LLMs show promise in argument assessment
Representative Image. Credit: ChatGPT

Artificial intelligence is increasingly being integrated into education, particularly in writing and critical thinking development. A recent study titled "Leveraging Small LLMs for Argument Mining in Education: Argument Component Identification, Classification, and Assessment" by Lucile Favero, Juan Antonio Pérez-Ortiz, Tanja Käser, and Nuria Oliver, explores the potential of small open-source Large Language Models (LLMs) in analyzing and assessing argumentative structures in student essays. This research, published by ELLIS Alicante, Universitat d'Alacant, and École Polytechnique Fédérale de Lausanne (EPFL), investigates how lightweight AI models can provide targeted feedback on argumentation skills while ensuring privacy and computational efficiency in educational settings.

Role of argument mining in Education

Writing structured and coherent essays is a fundamental skill for students, but many struggle to develop strong argumentation due to limited individualized feedback. Traditional assessment methods rely on educators manually reviewing essays, which is time-consuming and constrained by high student-to-teacher ratios. Argument mining, a subfield of natural language processing (NLP), aims to automate the detection and classification of argument components in text, thereby facilitating more effective writing feedback. The study focuses on three core tasks: segmenting essays into argument components, classifying those components by type, and assessing their quality.

Previous approaches to argument mining have primarily relied on large, complex deep learning models that are resource-intensive and challenging to implement in educational environments. This study addresses these limitations by exploring the viability of small, open-source LLMs that can run locally on student devices, ensuring accessibility and data privacy while maintaining high performance.

Evaluating small LLMs in argument mining

The researchers tested three small, open-source LLMs - Qwen 2.5 7B, Llama 3.1 8B, and Gemma 2 9B - using two primary machine learning techniques: few-shot prompting and fine-tuning. The models were evaluated on the "Feedback Prize - Predicting Effective Arguments" dataset, a large corpus of essays written by students in grades 6-12.

The first task, argument segmentation, involved dividing an essay into distinct argumentative components. The study found that fine-tuned small LLMs outperformed traditional methods, such as Longformer-based models, in accurately segmenting argument structures. In the second task, argument classification, models were trained to categorize argument components into types such as claim, evidence, counterclaim, and rebuttal. Again, fine-tuning significantly improved classification accuracy compared to previous baselines. Lastly, argument quality assessment, which required evaluating the persuasiveness and clarity of each argument, showed mixed results. Few-shot prompting performed comparably to baseline models, but fine-tuned models exhibited stronger consistency in grading.

Overall, the results highlight that small, fine-tuned LLMs can effectively analyze and assess argument structures while requiring significantly less computational power than their larger counterparts.

Implications for educational technology

The findings of this study have significant implications for integrating AI-driven feedback into educational technology. By leveraging small, accessible LLMs, schools and educators can deploy AI-assisted writing support without relying on proprietary, cloud-based models that raise concerns about data privacy and accessibility. Additionally, real-time, AI-generated feedback can help students refine their argumentative writing skills by identifying weaknesses in structure, coherence, and persuasiveness.

Furthermore, the research emphasizes the need for interdisciplinary collaboration between AI researchers, educators, and policymakers to ensure the ethical and effective implementation of AI in education. Issues such as algorithmic bias, fairness in grading, and the interpretability of AI-generated feedback must be addressed to maximize the benefits of argument mining tools in real-world learning environments.

Future of AI in writing instruction

While this study demonstrates the potential of small LLMs in argument mining, challenges remain. Future research should focus on improving the robustness of AI-driven argument assessment, particularly in evaluating nuanced aspects of writing such as rhetorical effectiveness and logical coherence. Expanding datasets to include more diverse writing styles and educational contexts will also enhance model generalizability.

The integration of AI into writing education represents a promising step toward personalized learning, allowing students to receive detailed, individualized feedback without overburdening educators. As small LLMs continue to advance, they have the potential to revolutionize how students develop critical thinking and argumentative writing skills, making high-quality education more accessible and efficient for learners worldwide.

  • FIRST PUBLISHED IN:
  • Devdiscourse
Give Feedback

Use this form for editorial or site feedback. We usually reply within 2 to 3 working days.

By submitting, you agree that we may use your email address to respond.