Writing Assessment in the AI LandscapeWriting Assessment

Evaluating Student Writing in the AI Era: Instructor Practices, Judgments, and Tools

Grace LiPooja VegesnaHan Zhang Mina Lee
University of Chicago
Ongoing project Resources for Instructors

This work is in progress. For feedback, please contact Grace Li at graceli458@uchicago.edu.

Generative AI has changed what instructors can infer from a finished piece of student writing. A polished submission may reflect a student's effort and growth, but it may also raise questions about how the writing was produced and whether AI use aligned with course expectations.

This research project synthesizes instructor interviews and tool-probe findings to describe how instructors move from an initial concern to contextual observation, reflection, and response. Rather than positioning AI detection as a standalone answer, the project centers instructor judgment, student writing process, transparency, and learning-focused action.

The resulting artifacts translate these findings into guidance for instructors and design considerations for tool builders, while attending to fairness, privacy, trust, and classroom climate.

Project Summary

This research project explores how instructors evaluate student writing in the AI landscape: the process that they follow and the tools and resources they use to support their process. The project draws on interviews with University of Chicago instructors who teach writing-intensive courses, along with tool probes that surfaced how instructors interpret AI-detection, authorship, source-attribution, and writing-process tools.

The artifacts being developed from this research are intended to support two audiences: instructors navigating writing assessment in courses where AI tools are available, and designers building tools that aim to support instructors without replacing their judgment.

Tools and Features Covered

Tools directly covered in the research materials:

  • Pangram (AI detection scores, AI-assistance categories, segment-level analysis, confidence indicators)
  • Process Feedback (writing-process reports, editing time, revision patterns, copy-paste events)
  • DraftMarks (human-AI co-writing traces, AI-generated content, revision intensity, deleted or unused AI text)

Feature categories covered by the study include text-based AI detection, citation and source verification, writing-process records, and student-AI interaction and source attribution.

Summary of Findings

1Instructors described a common process when evaluating student submission for potential AI use: noticing something unexpected in a submission, gathering observations from multiple sources, reflecting on uncertainty and stakes, and choosing a response.
2AI-detection scores were most useful as contextual signals, especially at the extremes, but instructors resisted treating them as determinations because the scores are currently unreliable and are harder to contextualize beyond the number.
3More granular tool outputs, such as paragraph- or sentence-level flags, were useful when they provided concrete points for conversation, but they also introduced interpretation burdens and could conflict with instructor judgment.
4Tools that verify factual issues, such as citations or source accuracy, were received more positively because their outputs are easier to check than stylistic or pedagogical judgments.
5Writing-process and source-attribution tools can support richer conversations about how work was produced, but they also raise concerns about privacy, surveillance, transparency, and the risk of shifting from learning support to enforcement.

A Framework for Understanding Student Writing in the AI Landscape

The framework presents four steps that may arise as instructors evaluate student submissions for potential AI use.

It was synthesized from interviews with University of Chicago instructors who teach writing-intensive courses, with the goal of identifying how instructors evaluate student submissions in the context of potential AI use and the considerations in each step of the evaluation process.

Ways to Use the Framework

The framework is not intended to prescribe a single process, as an instructor's process is shaped by their own professional values and teaching contexts. Instead, instructors can use it to reflect on their current practices, make implicit evaluation decisions more explicit, and consider resources or tools that may be useful in their own classroom contexts.

Example uses include:

  • Reflect on their current evaluation process: Identify the steps, sources of information, and judgments they already rely on when questions about AI use arise.
  • Check assumptions before acting: Consider whether an initial concern may be shaped by expectations about a student, their writing ability, or what student writing "should" look like.
  • Plan conversations with students: Decide when and how to ask students about their drafting, revision, source use, or AI use in ways that provide additional context.
  • Revisit course and assignment design: Use recurring concerns to identify where AI policies, assignment expectations, process documentation, or instructional scaffolding could be made clearer.
  • Compare practices with colleagues: Use the framework as a shared vocabulary for discussing how instructors make judgments and respond to similar situations across courses.

Step 1: Notice What Raised the Question

Initial signals that further review may be needed.

As the instructor is reading through the student submission, certain aspects in the submission may raise the instructor's concern about potential AI use. These initial reactions provide a signal that further review is needed, but often do not provide enough information to make a decision.

The table below provides the different aspects that raise the instructor's concern and some questions that instructors ask themselves during the process.

AspectsQuestions for Consideration
  • Does the writing sound like this particular student? Like a human? Like AI?
  • Are there habits you typically associate with this student present here (e.g., recurring interests, stylistic choices, etc.)?
  • Are the phrases used in the writing reflective of typical student writing in the course (e.g., first-year student)?
  • Does the writing seem more polished or structurally smooth than what is typically expected?
  • Is the polish or smoothness even across the whole paper or does the register change at particular points? Where are those points?
  • Are there reasons other than AI that this writing might not match your expectation, such as the student's language background, writing training, or how they process and produce written work?
  • Does the writing contain concepts or vocabulary that is not covered in the course?
  • Does the writing make sophisticated claims without providing clear evidence or textual analysis to support the claim?
  • Does the writing rely on broad and generic claims rather than close engagement with the text?
  • What features in the writing influence your assessment of this student's writing level?
  • Does the paper feel generic or disconnected from the assignment prompt?
  • Does the paper engage with topics covered in the course materials or class discussions?

Step 2: Gather Observations

Sources of context instructors may consider.

If the first read raises concern, instructors often seek additional information that can help place that concern in context. The table below presents the different sources of information that instructors reference, the questions that they consider, and the limitations that instructors recognize. Alongside these sources, instructors also reflected on their own assumptions and potential biases about the student and how these might shape their interpretation of the submission.

SourceQuestions for ConsiderationLimitations

Instructors may start with re-reading the student submission to identify specific concerns that they have.

  • Are there certain passages that read differently from the rest?
  • Are there non-existent citations or use of text that was not covered in the course?
  • Is the structure of the essay too formulaic?
  • Are there inconsistencies in the levels of polish or analysis in the essay?
  • These features alone are often insufficient in determining AI use.

When available, instructors may compare the submission with prior assignments, drafts, in-class writing, or other writing samples.

  • Are there sudden changes in the student's voice, polish, structure, or sophistication of the writing?
  • Are there certain features that have stayed the same in the student's writing?
  • Improvements in a student's writing ability may reflect a student's growth and are not necessarily evidence of AI use.
  • Different assignments prompts may not be comparable.

When available, instructors may use tools like Google Docs version history or Process Feedback to understand a student's writing process.

  • How did the student draft and revise the submission?
  • Are there certain copy/paste patterns that seem concerning?
  • Does the revision history provide additional context for how certain sections were drafted?
  • Instructors may not feel comfortable using writing-process-based tools in their evaluation process.
  • Writing-process tools might not capture the full writing process, if the student writes using different documents or on paper.

When available, instructors may consider using their interactions with the student in class discussion or office hours.

  • How has the student demonstrated their understanding of the course materials?
  • Does the submission engage with course-specific concepts, readings, discussions, or terminology?
  • Is the level or type of analysis consistent with opportunities students have had to develop these skills in the course? In other courses (e.g., core course or sequences)?
  • Student's spoken participation or class engagement may not always align with their written ability.

When available, instructors may ask colleagues or the Office of College Community Standards (OCCS) for additional support when evaluating a student submission.

  • Did colleagues or OCCS provide the same or different interpretation of the submission?
  • What types of information did colleagues or OCCS use to contextualize their judgement?

When available, instructors also might use AI detection tools or writing-process tools to provide additional information for them to consider.

  • Colleagues or OCCS may interpret the same submission differently.
  • It is also important to consider and recognize the instructor's own bias in the grading process.
  • Different AI detection tools can generate different results for the same essay.

Step 3: Reflect Before Choosing a Response

Learning goals, certainty, and stakes.

Before choosing a response, instructors often considered these three aspects: learning goals, their certainty, and the stakes.

CategoryQuestions for Consideration
  • Could this student have produced this work?
  • Is the work demonstrating that the student is learning?
  • Given the context that was gathered, how many conflicting pieces of information were gathered?
  • How confident in the gathered information is the instructor in making a decision?
  • What aspects still remain uncertain?
  • Where in the quarter did the incident occur and is there enough time for the student to course correct?
  • Does the instructor have enough bandwidth and energy to pursue this response?
  • What are the potential consequences if the instructor made a mistake?
  • How might the instructor's response impact the classroom environment?
  • What intervention would be best in supporting the student's future learning?

Step 4: Respond

Possible actions based on context and policy.

Instructors described several possible responses, depending on the available context, course policy, and learning goals:

ResponseQuestions for ConsiderationLimitations

Grade the submission on its merits when uncertainty remains or the available observations do not support further action.

  • What would you have needed to see in order to be more confident and could the next assignment or a variation of it show you that?
  • Declining to act can create inconsistency across students in the same course, including for students who followed the policy.
  • The uncertainty can affect how the instructor might read the student's future submissions.

Direct Conversation: The instructor directly raises concerns about a student's submission with the student and invites them to discuss their work, writing process, or use of AI.

  • Which specific passages will you ask the student to explain? What would an acceptable/satisfying answer include?
  • What types of information would you present to the student?
  • How might you structure your conversation?

Indirect Conversation: The instructor asks more broadly about the student's drafting and writing process without directly raising concerns about potential AI use.

  • Which concepts would a student who wrote this be able to explain?
  • What types of questions would you ask the student to demonstrate their understanding of the topic? What would be an acceptable response?
  • Requiring additional instructor effort and emotional bandwidth to prepare for these conversations.
  • Students might also feel stress leading up to and during these conversations.
  • Direct conversations can introduce new uncertainties to the instructor. A student may deny, disclose partially, or give reasoning that the instructor has no way to verify.
  • For indirect conversations, speaking about a text is a different skill than writing one and may not be a clear indication of AI use during writing.

Apply the course policy through a grading adjustment when the available context supports treating the case as a course-policy issue. Instructors may also choose to leave comments on the student's paper denoting specific areas that raised their concern.

  • Point deductions
  • Automatic zero
  • Might not directly address the root cause, i.e., might not address the unmet learning goal.
  • Requires a course policy defining the response the instructor will follow.
  • Inconsistency in enforcing the course policy can lead to fairness issues.

Offer the student a chance to demonstrate their learning through a revised submission or a different assessment format, giving the student another opportunity.

  • Should the student revise and resubmit the same assignment? Different assignment? Oral Exam? Or a written closed-book exam?
  • How might a different assessment format complement or complicate the instructor's evaluation of the student's learning?
  • In what ways could the new task make the desired learning visible?
  • Additional grading and scheduling load for the instructor in order to design and implement.
  • Timed, oral or closed-book formats measure different outcomes than take-home writing does. Students may perform differently depending on the format.

Consult teaching colleagues, the Office of College Community Standards (OCCS), or the appropriate Dean of Students office, and follow institutional guidance when formal reporting is appropriate.

  • What are the specific questions or passages that are causing uncertainty that you are bringing to them?
  • In what way would a second opinion support your decision?
  • What would you do depending on what they say?
  • Formal processes can take longer than the remaining term, which requires substantial time and energy for instructors and students to prepare for.
  • Institutional standards of evidence may not match what the instructor finds persuasive.
  • Colleagues and OCCS may interpret the same submission differently which could introduce additional uncertainty rather than resolving the case.

Tools for Understanding Student Writing in the AI Landscape

This summary organizes tools by the kinds of information they provide. The table below identifies the tools directly covered in the research materials and related tools with overlapping features. The goal is not to rank or endorse tools, but to clarify what each tool family can show, where its outputs may be useful, and what risks instructors may want to consider before using it in a course.

Across our interviews with instructors who teach writing-intensive courses at the University of Chicago, instructors treated these tools as sources of context. They were most useful when they helped instructors understand how a submission may have been produced, identify concrete points for conversation, or reduce manual checking work. They were least useful when they appeared to make a judgment that instructors saw as pedagogical, contextual, or relational.

Tools Directly Covered in the Research Materials

Tool Features covered or relevant to this study Related tools with overlapping features How instructors might use this information
AI detection scores; AI-assistance categories; segment-level analysis; confidence indicators. Turnitin AI Writing Report; GPTZero; Copyleaks AI Detector. Use detector outputs as one contextual signal when deciding whether to gather more information, re-read specific passages, or prepare questions for a conversation with the student.
Writing-process reports; editing time; revision patterns; copy-paste events. Google Docs version history; Draftback; Grammarly Authorship. Use writing-process records to understand how a submission developed and to support reflective conversations about drafting, revision, source use, and effort.
Human-AI co-writing traces; AI-generated content; revision intensity; deleted or unused AI text. Grammarly Authorship. Use source-attribution or AI-interaction records to discuss what role AI played in the writing process, what the student accepted or changed, and how those choices relate to course expectations.

1. Text-Based AI Detection

Scores, highlights, AI-generated categories, phrase-level signals.

What it shows

  • Document-level scores
  • Paragraph- or sentence-level highlighting
  • AI-generated vs. AI-paraphrased categories
  • Word- or phrase-level signals
  • Confidence indicators

How it works

These tools analyze submitted text and estimate whether portions resemble writing produced or revised by generative AI. Different tools report the estimate at different levels: overall document, segment, paragraph, sentence, or phrase.

What it can help with

  • Document-level results can help instructors decide whether additional review is needed, especially for essays with very high or very low AI detection scores.
  • Paragraph- and sentence-level flags can provide concrete reference points for instructors to use in their conversations with students without relying on these indicators to signal AI use.

Where it falls short

  • Middle-range AI detection scores (e.g., 40-70% AI-generated) are difficult to interpret.
  • Sentence or word-level AI detection results can conflict with instructor judgment, especially when a sentence reflects a student's ordinary academic register, an attempt at sophisticated writing, or weak writing.
  • Word-level signals are especially fragile because vocabulary choices can reflect many things besides AI use.
  • Students can access AI detection tools themselves. It is possible for a student to run their essay through detection tools repeatedly until it scores 0% AI-generated.

Considerations for classroom use

  • Are students aware of the usage of the tools being used? Are they allowed to see the results?
  • What type of AI detection tool will you use? How did the company evaluate the reliability of the AI detection tool (e.g., false positives, false negatives)?
  • How will AI detection tools be used for student submissions? Will all student submissions be run through an AI detection tool or will only suspected student submissions be run through an AI detection tool?
  • What is the AI detection score threshold that will determine whether further action is required? Will this threshold be shared with students?
  • Does your departmental policy require a formal case for an AI detection score above a certain threshold?
  • What will happen if the tool shows a result you disagree with?

2. Citation and Source Verification

Source existence, citation accuracy, similarity reports.

Tools

What it shows

  • Checks for fabricated, inaccurate, or misattributed sources
  • Comparison against source databases
  • Citation or similarity reports

How it works

These tools compare references, quotations, or source claims against existing bibliographic records, web sources, or submitted-text databases.

What it can help with

  • Instructors responded positively to citation verification because its outputs are easier to verify.
  • A source that does not exist, an inaccurate page range, or a misattributed quotation can be verified and discussed with students to promote better citation practices.
  • The feature also reduces manual checking work required by instructors.

Where it falls short

  • Citation problems do not, on their own, explain how a text was produced.
  • They may reflect AI use, weak research practices, citation-management errors, or misunderstanding of the assignment.

Considerations for classroom use

  • What is the procedure for checking citations? Will you check every citation or only ones that look unusual?
  • To what extent do different citation issues raise concerns in the context of your course, and how serious is each type of issue?
  • What are the existing citation practices that are taught in class?

3. Writing-Process Records

Version history, replay, copy-paste events, revisions.

What it shows

  • Version history
  • Replay of document creation
  • Insertion and deletion timelines
  • Copy-paste history
  • Editing time
  • Revision patterns

How it works

These tools show how a document unfolded over time, either by replaying revision history or summarizing process traces such as time, pauses, revisions, and copy-paste events.

What it can help with

  • Process records can make an otherwise invisible writing process visible.
  • Copy-paste history was especially useful to instructors because it can show what entered the document all at once and sometimes where it came from.
  • Aggregated process summaries can help instructors understand effort, revision, and drafting behavior without watching every keystroke.

Where it falls short

  • Replay tools can be time-consuming to review and can feel invasive.
  • Instructors worried that visible monitoring could change how students write and damage classroom trust.
  • Process records are also incomplete: students can retype generated text, draft elsewhere, use another device, or move text between documents.

Considerations for classroom use

  • What will you do if some students prefer to draft on paper, in another application, or with assistive technology? Would requiring students to write within a given document disadvantage anyone in the course? Are there alternatives that students can use?
  • What will be your procedure for reviewing the process outcomes for each student? On an "as needed" basis? Or all students?
  • Will you communicate the types of writing-process tools you use and how you use them to your students? This includes the types of information captured from the tool and how you might use this information in your evaluation process.
  • How might writing-process-based insights support students' learning or structure conversations with students about their writing process?
  • Which of these features are you comfortable and uncomfortable using? Why? Writing-process reports can vary widely in how much detail they provide to you about the student's writing process.

4. Student-AI Interaction and Source Attribution

Human vs. AI authorship, captured prompts, AI-generated output, edit history.

What it shows

  • Whether a passage was authored by the human writer or generated by AI
  • The prompt and instruction used when a passage was AI-generated, when that exchange is captured
  • The complete text of the AI-generated output
  • Version snapshots marking when AI-generated text was inserted, fully removed, or substantially edited

How it works

DraftMarks tracks writer-AI interaction across a few common setups: writing and AI chat in separate windows (e.g., a Google Doc alongside a ChatGPT tab), an AI assistant built directly into the writing tool, or an ambient AI assistant added through a browser extension. It classifies text mainly by whether it can be mapped to the writer's own keystrokes: text typed by the writer is marked as human-authored, while text that can't be traced to keystrokes, including pasted content that didn't originate in the same app, is marked as AI-generated. When the AI chat happens in a separate window, the tool infers AI origin from paste and keystroke patterns rather than directly capturing the prompt exchanged with the chatbot.

What it can help with

  • Interaction records provide a record rather than a prediction, which instructors found useful.
  • Seeing that a passage was generated, pasted, revised, or typed can clarify the scope and role of AI assistance.
  • These records can also support conversations about acceptable use: what the student asked AI to do, what they accepted or changed, and whether they understood the resulting text.

Where it falls short

  • These tools only capture activity inside the tracked setup (the supported editor, extension, or paired chat window).
  • When writing and an AI chatbot are in separate windows, the tool infers AI authorship from paste and keystroke patterns rather than directly recording the prompt and response, so the captured "prompt" may be incomplete or approximate.
  • A student can use an AI chatbot or workflow the tool isn't set up to track, another device, another browser, a private window, or any other untracked path.
  • The record may be incomplete even when no one intended to evade it.

Considerations for classroom use

  • What type of support from AI does your course policy permit?
  • Does your policy distinguish AI-generated, AI-assisted, feedback from AI on student written work, and grammar or spelling assistance?
  • Will students declare their AI use themselves, will the tool record it, or both?
  • If collecting student self-declaration, will it be in the form of listing AI assistance? Or asking for a reflection assignment on how AI contributed to their process?
Classroom use

Reflective Questions for Classroom Use

Tool scope: What tool, if any, might be used, and how will students be told in advance?
Use pattern: Will the tool be used for all submissions, selected submissions, or only after an instructor has first reviewed the work?
Data handling: What data will the tool access, store, or share? Are these tools licensed by the university?
Student visibility: What will students be able to see about the tool's output? Will you share the results of AI detection tools or writing-process tools with students?
Interpretation: How will instructors explain that a score, flag, or process record is contextual information rather than a standalone conclusion?
Corroboration: What other information will be considered alongside the tool outputs? How will you balance the two different sources of information?
Conversation: How can the tool support a conversation about the student's writing process rather than shift the interaction into a punitive frame?