AGI Soon As Possible · Deep reads on AI & tech
Article

A 1,053-Student Bocconi RCT: ChatGPT Raised Graded Scores by 0.862 Points While Causal-Reasoning Training Widened Idea Diversity Instead

2026-08-29 · 9 min read

OpenAI Economic Research and Bocconi University researchers released a 2×2 randomized controlled trial run on 1,053 first-year undergraduates in August 2026. The paper reports that ChatGPT Edu access (GPT-4o) raised the five-point evaluation score by 0.862 points against an estimated control score of 2.09, while causal-reasoning training raised mechanism identification by about 0.55 standard deviations and falsification logic by about 0.85 standard deviations without moving the score at all. The two treatments moved different capabilities, and combining them delivered both effects. ASAP reads the real finding as a statement about the rubric rather than about the students.

The design randomized whole classrooms into four rooms

The researchers ran a preregistered 2×2 factorial RCT across 13 first-year classes in economics and management, economics and finance, and management at a leading European university. OpenAI's announcement names the institution as Bocconi University, and most of the paper's authors hold Bocconi affiliations. Randomization happened at the classroom level within program, crossing causal-reasoning training against a placebo game with ChatGPT Edu access against none. Participant counts were 249 control, 256 causal, 197 GPT, and 351 combined, with the combined cell deliberately oversampled to gain statistical power on the interaction.

The task was a real one. During class time students spent 45 minutes writing up to 180 words of consultant-style recommendations for increasing alumni awareness and usage of the university's merchandising shop, two established criteria in the marketing literature. Students were not told about the experiment in advance, so participation reflected ordinary class attendance. Evaluation ran along two tracks: 20 master's students served as raters, three per submission, scoring on a 1-5 Likert scale, and automated text analysis separately measured distance from recommendations written by three domain experts. Self-reported compliance among participants with GPT access was 90.5% overall, 88% in the GPT-only cell and 92% in the combined cell.

The design choice worth noting is that the causal-reasoning training was unrelated to AI. It taught the cognitive skill itself through an exercise built from a game, examples, questions, and feedback. That separation lets the study isolate the effect of handing someone a tool from the effect of training how they think, and then observe the two in combination. Most AI research randomizes tool access only and treats cognitive ability as an observational variable participants happen to arrive with.

Only ChatGPT moved the score

ChatGPT Edu access is the only one of the two treatments that improves evaluation performance, and with balance controls it raises the awareness-and-usage score by 0.862 points against an estimated control score of 2.09. On a five-point scale that moves a submission close to a full step up from just below the midpoint. Submissions from the GPT group also landed closer to the recommendations written by the three domain experts. Causal training alone had no effect or a slightly negative one, and the same held for the combined treatment.

To separate presentation from substance, the researchers ran post-double-selection lasso regressions. The GPT coefficient starts at 0.862 and falls to 0.490, then 0.448, then 0.362 as controls are layered in. Roughly half the original effect survives after controlling for text features, number and diversity of ideas, and causal-reasoning measures. On that basis the paper concludes that the LLM improved the substance of the solutions and not only their presentation.

Both halves of that number deserve weight. The decline from 0.862 to 0.362 is direct evidence that much of the GPT effect came from coherence and idea count, which is exactly the long-standing worry that graders reward polish. The residual 0.362 pushes the other way: after stripping every measurable text feature, something remained that novices did not produce on their own. Both readings hold only under this study's conditions, namely novices facing a well-defined problem on which LLMs are heavily trained.

The training changed what the rubric was not looking at

The axis causal training moves is reasoning quality, and relative to control it raises mechanism identification by about 0.55 standard deviations and falsification logic by about 0.85 standard deviations. Trained students wrote more clearly about why their ideas would work and under what conditions they would fail. Idea diversity moved too. On within-solution diversity, measuring the distance between ideas inside a single submission, the causal condition produced the strongest effect at roughly 0.5 standard deviations, and on between-solution diversity, measuring distance across participants' core ideas, causal submissions sat further apart. GPT had a positive but much smaller effect on within-solution diversity and essentially none on between-solution diversity.

GPT instead drove idea count. It raised the number of core ideas by more than half a standard deviation and improved logical coherence. On its own, however, it induced neither mechanism-based reasoning nor falsifiability. Put plainly, GPT produced more ideas pointing the same direction, more smoothly, while causal training produced ideas pointing in different directions.

That orthogonality is the structural finding. The education debate is usually posed as a choice: should students learn to think for themselves or learn to use AI. This data says the two capabilities are not competing but perpendicular. One governs the density and finish of an answer; the other governs the breadth of directions an answer can take. Choosing between them in curriculum design is a badly framed problem to begin with.

Better thinking failed to become a better score because of the rubric, not the student

The result most in need of interpretation is that causal training changed how students reasoned without raising their scores. The paper's decomposition explains why. Falsifiability and mechanism identification were negatively related to the evaluation score, and once both were included in the regression the causal treatment's coefficient turned positive. The diversity measures split the same way: within-solution diversity was positive, between-solution diversity negative. Evaluators rewarded a rich set of ideas inside one answer and penalized answers that departed from the typical solution space.

The implication points at the assessment instrument rather than the students. The rubric was built to measure only how well a submission addressed two standard marketing goals, awareness and usage. Judged on that, an unconventional idea reads as deviation rather than improvement. In the paper's own words, the rubric penalizes distance from the conventional answer. The diversity the training produced was real and measurable by text analysis, yet inside an assessment system that does not demand diversity it may as well not have existed.

Crucially, this holds independent of the LLM. GPT neither increased nor decreased diversity. The failure to reward originality was not created by AI; it predates AI, and AI merely made polished answers cheap enough that the failure became visible. Deciding to ban or permit AI use without touching the rubric leaves the actual problem untouched.

Students kept using the trained skill even with the tool in reach

For education policy the most useful result comes from the combined cell. Students who received both ChatGPT and causal training showed both effects. Their idea diversity matched the training-only group, and their scores and idea counts matched the GPT-only group. On coherent logic the interaction term ran about 0.48 standard deviations positive while causal alone reached only 0.15. On mechanism identification the interaction was small but positive at roughly 0.20. Evidence of seeking explanations and questioning assumptions was clearest in the combined group.

What that amounts to is counter-evidence against the substitution worry. The fear that a student with AI within reach will simply hand the task over rather than apply a trained skill was not supported here. The training's effect survived the availability of an LLM. The authors note their design does not trace the underlying cognitive process, but they read the pattern as participants exercising the trained skill rather than delegating.

The conclusion invites an oversimplification worth resisting, namely that combining the two multiplies the benefit. The data is more cautious. The combined group did not significantly beat GPT-only on score, nor training-only on diversity. Combining gets each ceiling, not a product of the two. Even so, the combined group gained across the widest range of measures, which means any design that picks only one of the two is leaving something on the table.

Conditions and limits when carrying this into other systems

What this study aims at is not the debate over whether AI use should be permitted but the design of assessment itself. Grading a finished artifact against a standard rubric is pervasive across university coursework, corporate application screening, and internship evaluations. This experiment shows that under those conditions novices with an LLM score higher, and that roughly half the gain comes from the presentation layer. A gate that used to sort candidates by polish now sorts them by tool access, and preserving any discriminating power requires changing what is graded.

The paper's prescription is explicit: without a system that demands and evaluates diversity, the breadth of thinking that causal training produces never surfaces as performance. In practice that means asking for more than one final answer, requiring the alternatives considered and discarded, requiring a statement of the conditions under which the proposal fails, and assigning points to mechanism and evidence. Those are precisely the dimensions on which the trained group led here.

The limits on generalization should be stated as clearly. The model was GPT-4o, a prior generation by current standards, and effect sizes may differ with newer models. The task was a single 45-minute marketing case, not a semester of learning outcomes. Raters were 20 master's students rather than professional graders, and the expert baseline was three sets of recommendations. The population was first-year economics, management, and finance students at one university, and nothing in the paper licenses transfer to other disciplines or to entry-level professional roles. The authors also withhold the university name and the pre-analysis plan registration number from the text to preserve anonymity.

Source: OpenAI Economic Research announcement "Better answers, broader thinking" (August 27, 2026) and the paper "Training novices to think, or giving them LLMs? Evidence from an RCT" (August 2026, Bocconi University, OpenAI, Duke University, UC Berkeley), compiled by ASAP.

ASAP — AGI Soon As Possible

AI & tech,
read in depth

Beyond the headlines — into the context and the structure

AGI Soon As Possible · asapai.co.kr

← All posts