OpenAI has published the results of a randomized education experiment run with researchers at Bocconi University (Università Bocconi) in Italy[1]. Working with more than 1,000 first-year undergraduates, the study separated the effect of ChatGPT access from the effect of critical-thinking training. The two turned out to work in clearly different ways, and to complement each other.
A four-group experimental design
The setting was a real-world business case tackled by more than 1,000 first-year undergraduates at Bocconi University[1]. Students developed marketing recommendations for the university's merchandise store, and were randomly assigned by class period to one of four groups: ChatGPT access (GPT-4o), training in causal reasoning, both, or neither.
The causal reasoning training targets the ability to link cause and effect and to explain why a proposed solution might work, or when it might fail. It had nothing to do with AI, teaching the concepts through a game-based exercise, worked examples, questions, and feedback[1].
Submissions were scored by trained human graders on a five-point rubric. Separately, automated text analysis measured the number and variety of ideas in each submission, signs of causal reasoning, and similarity to recommendations written by three experts[1].
ChatGPT access lifted scores by almost a full point
Students with ChatGPT access scored almost a full point higher on the five-point scale[1]. Their answers contained more ideas, followed clearer logic, and resembled expert recommendations more closely. In short, AI helped novices produce work that looked considerably more professional.
The researchers note that students were not simply handing the assignment over to ChatGPT. Deciding what to ask, evaluating the responses, and choosing what made it into the final submission all remained the student's job[1].
The training effect never showed up in the rubric
The critical-thinking exercise produced the more surprising result. Students who completed it explained more clearly why their ideas might work and when they might fail, yet their rubric scores did not improve[1].
The explanation lies in how the rubric was built. It measured only how well a recommendation addressed two standard marketing goals: increasing awareness and use of the university store[1]. Automated text analysis revealed a different effect. Students who completed the exercise generated a wider range of ideas that were more distinct from what their peers produced. A conventional rubric can reward a clean, well-structured answer while overlooking an idea nobody else came up with.
The group that received both improved most broadly
Students who received both ChatGPT access and the critical-thinking exercise showed the benefits of each[1]. Their idea variety matched the exercise-only group, while their rubric scores and idea counts matched the ChatGPT-only group. Their work also showed stronger logical coherence and more evidence of searching for explanations and questioning assumptions. Across the full set of measures, this group posted the widest range of gains.
The education debate is often framed as a choice between teaching students to think for themselves and teaching them to use AI. These results point somewhere else: both matter, in different ways.
The measuring stick itself is under question
The heavier implication concerns how schools assess work. Once AI makes it easy to produce polished, expert-looking output, the final answer alone says less about what a student actually understands[1]. OpenAI's reading is that originality, reasoning, and evidence of considering multiple approaches need to become part of what gets rewarded.
Bocconi University signed an agreement with OpenAI in May 2025, becoming the first university in Italy to offer AI tools to all 17,000 of its students, faculty, and staff[2]. In March this year OpenAI also introduced its Learning Outcomes Measurement Suite for tracking learning outcomes over time, and named Bocconi among the institutions hosting research at the intersection of learning and labor[3]. This experiment sits within that broader program.
Summary
The value of this experiment lies in using a randomized design to separate two effects of a different nature: ChatGPT access raised the quality and logical coherence of answers, while critical-thinking training broadened the range of ideas. That said, the study covers a single assignment by first-year undergraduates at one university over a limited period. Whether the findings transfer to other educational settings will depend on replication.
Source: https://openai.com/index/what-students-gain-from-chatgpt-critical-thinking-training
Source: https://www.unibocconi.it/en/news/bocconi-and-openai-join-forces-transform-ai-social-sciences
Source: https://openai.com/index/understanding-ai-and-learning-outcomes/
