Table of Contents
Fetching ...

Impact of AI Tools on Learning Outcomes: Decreasing Knowledge and Over-Reliance

Márton Benedek, Balázs R. Sziklai

TL;DR

Problem: The paper addresses whether unrestricted AI tool use affects motivation and genuine knowledge in higher education. Approach: a randomized, field experiment in a Corvinus OR course compared AI-permitted and offline learning with a compensation mechanism, followed by a last-minute merger of groups due to student protests. Key findings: offline knowledge was not clearly superior; paper-test results indicate knowledge near random guessing and an estimated knowledge decline of $20$–$40$ percentage points; analyses using SRD and linear regression show weak alignment between test scores and knowledge, while AI-detection reveals near-universal reliance. Significance: results call for cautious AI integration in education, emphasize perceptual and incentive effects, and illustrate methodological challenges in field experiments dealing with sensitive technologies.

Abstract

Students at all levels of education are increasingly relying on generative artificial intelligence (AI) tools to complete assignments and achieve higher exam scores. However, it remains unclear how this reliance affects their motivation, their genuine understanding of the material, and the extent to which it substitutes for the process of knowledge acquisition. To investigate the impact of generative AI on learning outcomes, an experiment was conducted at Corvinus University of Budapest. In an operations research class, students were randomly assigned into two groups: one was permitted to use AI tools during classes and examinations, while the other was not. To ensure fairness, a compensation mechanism was introduced: students in the lower-performing group received point adjustments until the average performance of the two groups was equalized. Despite the organizers' best efforts to explain the design and to create equal opportunities for all participants, many students perceived the experiment as a major disruption. Although the experiment was approved by every relevant university authority -- including the Ethics Board, the Head of Department, the Program Director, and the Student Council -- students escalated their concerns to the media and eventually to the State Secretary for Higher Education of Hungary. As a result, the experiment had to be substantially revised before completion: on the final exam the test group was merged with the control group. Still, the data allowed us to draw decisive conclusions regarding the students' learning habits. Uncontrolled use of AI tools leads to disengaged students and low understanding of material. The extreme reactions of the students proved even more revealing than the data collected: generative AI tools have already become indispensable for students, raising fundamental questions about the validity of their learning process.

Impact of AI Tools on Learning Outcomes: Decreasing Knowledge and Over-Reliance

TL;DR

Problem: The paper addresses whether unrestricted AI tool use affects motivation and genuine knowledge in higher education. Approach: a randomized, field experiment in a Corvinus OR course compared AI-permitted and offline learning with a compensation mechanism, followed by a last-minute merger of groups due to student protests. Key findings: offline knowledge was not clearly superior; paper-test results indicate knowledge near random guessing and an estimated knowledge decline of percentage points; analyses using SRD and linear regression show weak alignment between test scores and knowledge, while AI-detection reveals near-universal reliance. Significance: results call for cautious AI integration in education, emphasize perceptual and incentive effects, and illustrate methodological challenges in field experiments dealing with sensitive technologies.

Abstract

Students at all levels of education are increasingly relying on generative artificial intelligence (AI) tools to complete assignments and achieve higher exam scores. However, it remains unclear how this reliance affects their motivation, their genuine understanding of the material, and the extent to which it substitutes for the process of knowledge acquisition. To investigate the impact of generative AI on learning outcomes, an experiment was conducted at Corvinus University of Budapest. In an operations research class, students were randomly assigned into two groups: one was permitted to use AI tools during classes and examinations, while the other was not. To ensure fairness, a compensation mechanism was introduced: students in the lower-performing group received point adjustments until the average performance of the two groups was equalized. Despite the organizers' best efforts to explain the design and to create equal opportunities for all participants, many students perceived the experiment as a major disruption. Although the experiment was approved by every relevant university authority -- including the Ethics Board, the Head of Department, the Program Director, and the Student Council -- students escalated their concerns to the media and eventually to the State Secretary for Higher Education of Hungary. As a result, the experiment had to be substantially revised before completion: on the final exam the test group was merged with the control group. Still, the data allowed us to draw decisive conclusions regarding the students' learning habits. Uncontrolled use of AI tools leads to disengaged students and low understanding of material. The extreme reactions of the students proved even more revealing than the data collected: generative AI tools have already become indispensable for students, raising fundamental questions about the validity of their learning process.
Paper Structure (10 sections, 5 figures, 2 tables)

This paper contains 10 sections, 5 figures, 2 tables.

Figures (5)

  • Figure 1: Students' grasp of the material as measured by the paper test. Neither group demonstrates knowledge levels significantly different from random guessing. Left: How paper test scores compare to knowledge levels modeled by binomial distribution. Right: Likelihood that, given the test scores, students knew $x$ questions correctly and guessed the remaining $50-x$.
  • Figure 2: Ranking distance between the paper test and other test scores compared to distances of random rankings measured in SRD. XX1 and XX19 denotes the 0.05 and 0.95 significance threshold respectively.
  • Figure 3: Final exam scores by true-or-false results. $R^2$ values indicate almost no relationship, although the offline group shows a slight visual trend.
  • Figure 4: Final exam AI content assessed by an AI detection tool
  • Figure 5: Example graph. Edge weights indicate construction costs.