OpenAI just published results from a fascinating randomized experiment that should reshape how we think about AI in education. Working with researchers at Bocconi University, they ran a real-world study on more than 1,000 first-year undergraduates tackling an actual business assignment: developing marketing recommendations for the university's merchandise store.
The setup was elegant. Students were randomly assigned to four groups: ChatGPT (GPT-4o) access, critical-thinking training (specifically causal reasoning exercises), both, or neither. Then trained human graders evaluated their work on a five-point rubric, while automated text analysis measured idea variety, logical coherence, and similarity to expert submissions.
The punchline? ChatGPT and critical-thinking training each improved student performance—but in completely different, complementary ways.
AI Closed the Expertise Gap
Students with ChatGPT access scored almost a full point higher on the five-point scale. Their submissions included more ideas, followed clearer logic, and looked more similar to what experts produced. In other words, AI helped novices write like professionals.
This wasn't copy-paste laziness. Students still had to formulate good prompts, evaluate ChatGPT's responses, and make editorial decisions about what belonged in their final submission. The AI acted as a force multiplier for quality and coherence.
The text analysis backed this up quantitatively: more ideas per submission, tighter logical structure, expert-adjacent recommendations. If you're optimizing purely for rubric scores and professional polish, giving students ChatGPT access is a no-brainer.
Critical Thinking Delivered Something the Rubric Missed
Here's where it gets interesting. The critical-thinking exercise—an unrelated training module on causal reasoning using games, examples, and feedback—didn't improve rubric scores at all.
Wait, what?
Students who completed the exercise explained more clearly why their ideas might work and when they might fail. But the traditional rubric only measured how well recommendations addressed two standard marketing goals: increasing awareness and use of the university store. It wasn't designed to reward originality or causal reasoning depth.
The automated text analysis revealed what human graders missed. Students who completed the critical-thinking exercise produced a wider range of ideas that were more distinct compared to their peers. They questioned assumptions, explored edge cases, and generated recommendations that didn't cluster around the same obvious solutions.
This is the insight that makes the whole study click: traditional rubrics reward polish and correctness, but they're blind to originality and reasoning diversity.
The Combo Effect: Best of Both Worlds
Students who received both ChatGPT access and the critical-thinking training showed benefits from each intervention.
Their idea variety matched students who only did the exercise. Their rubric scores and idea volume matched students with only ChatGPT access. But crucially, their work showed stronger logical coherence and more evidence of questioning assumptions and seeking explanations.
This group gained across the widest range of measures. They produced polished, expert-like submissions with the added bonus of original thinking and causal reasoning depth.
The randomized design is what makes this finding credible. By isolating each intervention and testing the combination, researchers could separate correlation from causation. It's not just that better students do better things—the interventions themselves produced measurable, distinct effects.
The Rubric Problem
The friction between what the rubric measured and what the text analysis revealed points to a structural challenge facing educators everywhere.
If AI can help students produce polished, expert-like work with minimal subject-matter expertise, then grading final outputs tells you less about what students actually understand. A well-structured answer might be the result of deep comprehension or skillful prompt engineering—and rubrics designed for the pre-AI era can't distinguish between them.
Many educators are already grappling with this. The study frames it as part of a multi-decade pattern: technology evolves, pedagogical assessment adapts in response. Calculators changed math education. Search engines changed research assignments. AI is forcing the next evolution.
The researchers argue that assignments and evaluations need to shift toward rewarding originality, reasoning processes, and consideration of multiple approaches—not just the most conventional or polished final answer.
What This Means for AI-Era Education
The framing that emerges from this study is refreshingly non-binary. The education debate often devolves into "should students think for themselves or use AI?" as if those are mutually exclusive.
This experiment shows both matter, in different ways:
- AI access improves quality, coherence, idea volume, and professional polish
- Critical-thinking training increases originality, idea diversity, and causal reasoning depth
- Together they're complementary, not competing
The "either/or" framing is a trap. Students need both capabilities.
What's particularly compelling is that the critical-thinking intervention had nothing to do with AI. It was a standalone exercise in causal reasoning. The implication: general-purpose thinking skills layer productively on top of AI tool access. You don't need to teach "AI literacy" as a separate domain—you need to teach reasoning, and let students apply it to whatever tools they have.
Open Questions
The study leaves some tantalizing threads dangling:
- How durable are these effects? Does ChatGPT access atrophy students' ability to generate ideas independently over time, or does it scaffold skill development?
- What happens when the rubric does reward originality? Would ChatGPT access still dominate, or would the critical-thinking intervention show stronger effects?
- Would similar results hold for other subjects (STEM, humanities, technical writing)? Marketing strategy feels like a domain where both polish and originality matter—what about fields where one clearly dominates?
The researchers acknowledge the rubric limitation explicitly, which is refreshing. Most education studies optimize for whatever metric they can measure, then declare victory. This one says "our rubric missed something important, and here's the text analysis that caught it."
The Takeaway
AI helped students make their answers better. Critical-thinking training helped make their ideas broader. The two play complementary roles.
For educators, the message is clear: if you're still grading purely on final-output quality, you're measuring what AI does well and ignoring what humans uniquely contribute. Assignments need to evolve to capture reasoning processes, originality, and the ability to question assumptions—not just polished final answers.
For AI builders, there's a design implication too. Tools that only optimize for output quality might accidentally train users to outsource thinking entirely. The sweet spot is AI that amplifies idea generation and polish while still requiring users to exercise judgment, evaluate tradeoffs, and defend their choices.
This study is a rare example of rigorous, randomized experimental design applied to a real-world educational context. It's the kind of evidence base we need as AI reshapes learning at scale. More of this, please.