Sources & Case Studies · Published investigation
A community interaction pattern, tested against bounded educational evidence.
A Tutor Prompt Is More Than a Refusal
A student-designed prompt replaces assignment-ready prose with explanation, questions, and another attempt. Research supports parts of that learning loop—but not universal refusal, guaranteed learning, or a new requirement for writing tools.
Source Signal
A prompt that asks the model to stop finishing the assignment.
A post in r/PromptEngineering describes a familiar sequence from the poster’s view as an undergraduate: paste an essay question, receive finished prose, reword it, and submit it. That is an observation from one community member, not a measured claim about how students generally use AI.
The alternative prompt changes the assignment given to the model. It prohibits paragraphs that can be pasted into the submission. It asks for one simple explanation, then asks the learner to explain the idea back. When the learner misses something, the model should ask a question that exposes the gap instead of immediately supplying the correction. The exchange continues until the learner can articulate the concept independently.
The post is useful because it describes an interaction design in unusually concrete terms. It is not a controlled evaluation. It does not establish improvements in grades, memory, writing quality, authorship, originality, or academic integrity. Engagement counts and comments do not strengthen those claims. [1]
Key Observation
Refusal is not the intervention. Retrieval, explanation, feedback, and another attempt are the intervention.
The important move is not refusal
Withholding an answer can create space for learning—or merely create a dead end.
“No paste-ready prose” is a boundary. On its own, it does not teach anything. A system can refuse help in a way that is arbitrary, inaccessible, or simply irritating. The educational claim begins only when the space after refusal contains a meaningful activity.
In the proposed loop, the learner retrieves an idea, explains a relationship, exposes a missing premise, receives bounded guidance, and tries again. The model’s role changes from replacing the task to organizing another encounter with it. The learner still has to decide what the concept means, whether the explanation holds together, and how to express it.
This distinction also prevents an easy moral conclusion. Finished prose is not inherently harmful, and effort is not inherently educational. The question is what human activity the task is meant to exercise—and whether the assistant preserves, supports, or replaces that activity.
Evidence
Parts of the loop have support. The complete prompt does not.
- Self-explanation: prompting learners to explain material can improve understanding in some instructional settings.
- Retrieval: attempting to reconstruct studied material can improve delayed retention compared with repeated study under some conditions.
- Supported challenge: prior knowledge, task complexity, cognitive load, guidance, and feedback help determine whether additional effort is productive.
- Emerging AI evidence: one recent randomized AI-access experiment reports different persistence patterns among post-treatment usage groups, but it did not randomize tutor-like versus generator-like interfaces.
What self-explanation research supports
Explaining can make a learner’s model available for revision.
In a foundational study, Michelene Chi and colleagues examined 24 eighth-grade students learning from a science text. Students who generated more effective self-explanations developed stronger understanding of the material. The setting was controlled, the sample was small, and the activity was not an AI-writing task. [2]
A later meta-analysis by Kiran Bisra and colleagues found a positive average effect for induced self-explanation across varied studies. That broader evidence makes self-explanation a legitimate instructional mechanism rather than a clever feature invented by one Reddit prompt. It also shows why implementation matters: prompts, comparison activities, subject matter, and outcome measures differ across studies. [3]
Self-explanation is useful partly because it makes the learner do more than recognize fluent material. An attempted explanation reveals what the learner connects, omits, or misunderstands. But asking for an explanation does not guarantee a good explanation, accurate feedback, or learning. An LLM that invites a learner to explain something does not automatically reproduce the conditions or results of these studies.
What retrieval practice adds
Receiving fluent material and reconstructing it are different cognitive events.
Henry Roediger and Jeffrey Karpicke found that retrieving studied prose improved delayed retention compared with repeated study, even though repeated study could produce greater confidence and stronger immediate performance. Their work supports a narrow but important distinction: an attempt to bring an idea back can strengthen later access to it. It does not prove that an AI should refuse to answer an essay request. [4]
John Dunlosky and colleagues later rated practice testing as a high-utility learning technique while emphasizing variation in materials, learners, implementation, and feedback. [5] A 2025 systematic and meta-analytic review by Ariel de Oliveira Gonçalves, Bruno Felipe Barbosa Muniz, and Antônio Jaeger further narrows any simple hierarchy. Retrieval practice held a small overall advantage over elaborative encoding, but the comparison changed substantially with feedback; without feedback, elaborative encoding could perform better. [6]
The comparison activity matters. Feedback matters. Retention is not writing quality, and recalling studied material is not the same as composing an original argument. Retrieval research makes the learner’s attempt worth considering; it does not turn refusal into a universal interface rule.
Friction can teach—or merely obstruct
The useful distinction is supported challenge, not more difficulty.
“Friction is good” is too blunt to guide an interface. Jason Lodge and colleagues describe learning difficulty and confusion as potentially productive or unproductive depending on prior knowledge, self-regulation, task design, and the timing of guidance and support. [7] Ouhao Chen, Juan C. Castro-Alonso, Fred Paas, and John Sweller similarly warn that added difficulty can become undesirable when complex material already places heavy demands on working memory. [8]
Obstructive friction includes unclear instructions, repeated failure without useful feedback, arbitrary withholding, and interface cost unrelated to the intended learning. Supported challenge asks for a relevant idea, a relationship, a choice between interpretations, or a revision—and provides a path forward when the learner cannot yet supply it.
That boundary matters most where prior knowledge or access cannot be assumed. A novice may need a worked example before self-explanation becomes useful. A multilingual learner may understand the concept while needing language support to express it. A disabled learner may require direct composition assistance, alternative input, reduced memory demand, or stronger scaffolding. Requiring everyone to struggle in the same way can preserve neither learning nor agency. Productive effort must be relevant, supported, and accessible.
Augmentation and automation are behaviors
The same AI access can support explanation or produce paste-ready text.
A July 2026 working paper by Zara Contractor and Germán Reyes provides emerging evidence about that distinction. The researchers randomly assigned 211 Middlebury College undergraduates to AI-allowed or AI-forbidden conditions during a proctored learning session. Students were assigned blockchain, carbon capture, or CRISPR material—topics on which they showed limited measured baseline knowledge—and wrote an analytical essay. Students given AI access could use any generative-AI tool and received a logged-in GPT-4o account. Immediate unaided knowledge scores rose by 6.7 percentage points, or 0.27 standard deviations. Among the 204 students who returned, the delayed effect one week later was 5.1 points, also 0.27 standard deviations. [9]
The experiment randomized access to AI. It did not randomize a tutor interface against a generator interface. After treatment, an LLM classified monitored ChatGPT conversations as augmentation, automation, mixed, or other. Augmentation covered explanation, clarification, feedback, and checking the student’s work. Automation covered generated essay text, paragraphs, or paste-ready material. Mixed users could enter both student-level subgroups, and each subgroup was compared separately with the full control group. The paper does not state that this classification was preregistered.
The subgroup pattern is suggestive, not definitive. The augmentation subgroup’s delayed-test estimate was positive but not conventionally statistically significant. Automation users showed a short-run essay-quality advantage that disappeared when AI was removed. That does not mean automation produced no learning: its delayed knowledge estimate remained positive but imprecise. The study supports further investigation of how people use the same tool. It does not establish a causal advantage for a tutor-only interface.
An earlier 2025 conference abstract reported a different retention conclusion. This article uses only the July 2026 arXiv and IZA paper for its evidence claims; the earlier abstract is version history, not a source to merge into the later results.
Authorship preservation is an editorial inference
The evidence measures learning activities, not authorship.
Learning, originality, legal authorship, intellectual participation, and visible decision-making are related questions, but they are not interchangeable. None of the studies cited here establishes that a tutor prompt preserves authorship. The Reddit post does not measure it, and the Contractor-Reyes experiment evaluates knowledge and essay outcomes rather than legal or philosophical ownership.
The narrower editorial inference is that asking a user to retrieve, explain, and choose can leave more of that person’s reasoning visible in the process. Their attempted explanation supplies material the assistant can question. Their choice determines which interpretation proceeds. Their revision shows whether the feedback changed the claim. This is preserved intellectual participation, not guaranteed originality.
That distinction keeps responsibility visible. A user who makes the claim should be able to examine its premise, degree of certainty, and consequences. But visible participation is not a purity test, and direct assistance does not erase a person’s contribution by definition.
Question for Later Testing
Could an optional learning-oriented workflow ask for an explanation or choice before offering a full revision?
That is a research question, not a roadmap item. Answering it would require a product specification, prototype, user testing, accessibility review, outcome measurement, and compatibility analysis with the current deterministic architecture.
The test would also need an escape route. A learning-oriented interaction should not trap a user who needs direct language assistance, and it should not confuse a model’s judgment with assessed mastery. The relevant comparison is not “help versus no help.” It is whether different forms and sequences of help preserve the activity a particular task is intended to exercise.
Two paths through the same request
The interface changes which activity remains visible.
Two interaction patterns. The second preserves more visible learner activity, but its educational value still depends on the task, prior knowledge, guidance, accessibility, factual quality, and feedback.
Direct generation
- Prompt
- Finished prose
- Submission
Learning-oriented exchange
- Prompt
- Explanation attempt
- Question about a gap
- Revised explanation
- Learner articulates the current understanding without a supplied final paragraph
Closing principle
A useful assistant should not maximize either output or obstruction.
It should preserve the human activity the task is meant to exercise. Sometimes that means producing a finished document. In a learning-oriented task, it may mean explaining less, asking better questions, and waiting for another attempt.
Continue the Record
This source-led investigation belongs to a broader account.
-
Editor’s Desk
The Record Behind Product Transparency
How this investigation expands the record from engine behavior into capability, limitation, and expectation.
Related Principles
- Evidence Before Narrative
- Editorial Transparency
Sources and notes
The evidence trail.
- r/PromptEngineering. “Stop letting ChatGPT be an ai writing tool for your essays. Here’s the prompt I give people instead.” Reddit, July 21, 2026. Community source signal; not an effectiveness study.
- Chi, Michelene T. H.; de Leeuw, Nicholas; Chiu, Mei-Hung; and LaVancher, Christian. “Eliciting Self-Explanations Improves Understanding.” Cognitive Science, 18(3), 439–477, 1994. DOI: 10.1207/s15516709cog1803_3.
- Bisra, Kiran; Liu, Qing; Nesbit, John C.; Salimi, Farimah; and Winne, Philip H. “Inducing Self-Explanation: a Meta-Analysis.” Educational Psychology Review, 30, 703–725, 2018. DOI: 10.1007/s10648-018-9434-x.
- Roediger, Henry L., III, and Karpicke, Jeffrey D. “Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention.” Psychological Science, 17(3), 249–255, 2006. DOI: 10.1111/j.1467-9280.2006.01693.x.
- Dunlosky, John; Rawson, Katherine A.; Marsh, Elizabeth J.; Nathan, Mitchell J.; and Willingham, Daniel T. “Improving Students’ Learning With Effective Learning Techniques: Promising Directions From Cognitive and Educational Psychology.” Psychological Science in the Public Interest, 14(1), 4–58, 2013. DOI: 10.1177/1529100612453266.
- Gonçalves, Ariel de Oliveira; Muniz, Bruno Felipe Barbosa; and Jaeger, Antônio. “Retrieval Practice Versus Elaborative Encoding: A Systematic and Meta-analytic Review.” Educational Psychology Review, 37, Article 100, 2025. Published October 30, 2025. DOI: 10.1007/s10648-025-10076-6.
- Lodge, Jason M.; Kennedy, Gregor; Lockyer, Lori; Arguel, Amael; and Pachman, Mariya. “Understanding Difficulties and Resulting Confusion in Learning: An Integrative Review.” Frontiers in Education, 3, Article 49, 2018. DOI: 10.3389/feduc.2018.00049.
- Chen, Ouhao; Castro-Alonso, Juan C.; Paas, Fred; and Sweller, John. “Undesirable Difficulty Effects in the Learning of High-Element Interactivity Materials.” Frontiers in Psychology, 9, Article 1483, 2018. DOI: 10.3389/fpsyg.2018.01483.
- Contractor, Zara, and Reyes, Germán. “Experimental Evidence on the Learning Impact of Generative AI.” arXiv:2607.08849v1, submitted July 9, 2026; IZA Discussion Paper No. 18792, July 2026.