Teacher Professional Development
Real data. No title yet. Look first, then we'll talk.
The Setup
54 participants, ages 18–39, from MIT, Wellesley, Harvard, Tufts, and Northeastern. Four sessions over ~four months. EEG measured brain activity throughout. Same eight questions asked after every session.
Task: write a 20-minute essay on an SAT-style prompt (loyalty, happiness, courage, etc.). Essays were analyzed with n-gram extraction, semantic embedding, and ontology mapping.
Three groups. Same condition for Sessions 1–3.
Group 1
ChatGPT was the only resource allowed. No other websites or tools.
Group 2
Any website except ChatGPT or any LLM. Most used Google with "-ai" to filter AI results.
Group 3
No internet. No tools. Memory and prior knowledge only.
Sessions 1 Through 3
The same questions after every essay. Select a session to see how each group responded.
Could quote their essay?
Perceived ownership?
Satisfied with essay?
0 of 18 LLM participants produced a correct quote from their own essay — immediately after writing it. Every Search Engine and Brain Only participant could. 3 of 18 LLM participants reported zero ownership. The LLM group was the most satisfied.
Could quote their essay?
Perceived ownership?
Satisfied with essay?
Participants knew the quoting question was coming — LLM recall improved, but 4 of 18 still couldn't produce a correct quote. Brain Only hit perfect recall (18/18). No zero-ownership reports this session, but LLM ownership remained the most fragmented.
Could quote their essay?
Perceived ownership?
Satisfied with essay?
Search Engine and Brain Only: perfect recall (18/18). 5 LLM participants still couldn't quote their essay. 1 reported zero ownership. Brain Only: all 18 reported full ownership. EEG showed LLM alpha-band engagement significantly lower than both other groups.
What the writing itself showed — N-gram analysis
N-gram analysis extracted the most common 2–5 word sequences per group. Patterns were distinct enough to identify group membership from the writing alone.
LLM — dominant phrases
teach children choose career personal successHigh consistency — shared phrases and concept maps across participants.
Search Engine — dominant phrases
homeless person moral obligation give usResearch woven into personal positions. More individual variation than LLM.
Brain Only — dominant phrases
true happiness benefit others think twice freedom libertyLeast overlap with other groups. Almost no shared conceptual edges across essays.
Two English teachers scored all essays blind
"Essays sounded academic but failed to give personal insights. We perceived these essays as 'soulless' — empty with regard to content, lacking personal nuances. We valued individuality and creativity over objective 'perfection.'"
AI essays scored higher on grammar. Lower on uniqueness and content. An AI judge rated nearly everything 4/5 — human teachers strongly disagreed on uniqueness. Human readers saw what the rubric missed.
Session 4 — The Tool Swap
Switched to the opposite condition. 18 of 54 returned (n=9 per group).
LLM group → no tools. Brain Only → ChatGPT. Each wrote on topics they had already covered in Sessions 1–3.
LLM group → now Brain Only (n=9)
Brain Only group → now using ChatGPT (n=9)
Session 4 results — visualized
Could quote their Session 4 essay?
Reported full ownership?
EEG finding: When the LLM group switched to unaided writing, their brain engagement did not recover. Alpha-band connectivity remained suppressed — the researchers call this accumulated cognitive debt that does not reverse in a single session.
The Brain Only group — now using ChatGPT — retained nearly everything: topic recognition, recall, ownership. Three months of unaided thinking protected them even when AI did some of the work. The LLM group, given the same task without tools, still couldn't remember or own most of what they wrote.
The source
MIT Media Lab — arXiv, June 2025 — 54 participants, 4 sessions over ~4 months
Limitations: Preprint, not yet peer-reviewed. Session 4 n=9 per group. All participants from elite Boston universities. EEG methodology drew scrutiny in published commentary. Self-report and n-gram data are reported directly from the study.