Tutor Reflection

A redesigned post-session reflection workflow that replaces open-ended writing with structured, context-aware prompts, turning a rushed compliance step into an intuitive experience researchers and supervisors can collect better data from. It was built for PLUSβ€”a CMU edtech platform that pairs middle school math learners with tutors, combining AI technology with research-backed methods.

Skills

Workflow Redesign

AI Interaction Design

Content Design

TEAM

1 Product Designer

1 Design Lead

Timeline

3 months

ROLE

Product Designer

PROBLEM

Multiple sessions a week, and the same form after each one. The repetition got tedious, leading to copy-pasted answers and empty fields.

Each section of the form used a five-star rating scale, text-heavy multiple choice, and a free-response field. From September 2025–January 2026, 74% of session fields and 40% of student fields were left empty, and 148 submissions were copy-pasted from a prior session. Additionally, an ambiguous issue escalation pathway caused confusion for tutors with real concerns and no formal way to report them.

How can we transform a rushed compliance step into something tutors are willing to slow down for?

As a CMU-affiliated edtech platform, PLUS depends on this form as a source of qualitative data, feeding both academic research and supervisor oversight of student outcomes. The redesign had to solve two problems at once: give tutors a format they'd actually engage with, and give researchers and supervisors data specific enough to act on. It also needed to map onto PLUS' goal-setting framework around student effort, progress, and engagement.

RESEARCH

It wasn't that tutors had nothing to say. The form wasn't asking the right questions or giving the structure needed to encourage thoughtful reflection.

THE DATA HANDOFF

This redesign started from known failure points:

Nothing structured for an AI system to learn from. πŸ—ƒοΈ

74% of session fields and 40% of student fields were left empty, and 148 submissions were copy-pasted from a prior session, resulting in sparse, duplicate, and inconsistent input.

Generic reflection prompts produce unreliable data. πŸ”

The form asked the same question regardless of what actually happened. There was nothing in the prompt itself to anchor a specific answer, so tutors defaulted to reusing an answer or leaving it blank.

The escalation pathway was ambiguous. ⚠️

A single yes/no toggle gave no way to say what was actually wrong, so routine friction and real safety concerns looked identical to supervisors, making it hard to spot which students needed intervention.

10–15 min per submission still wasn't producing usable data. ⏱️

Tutors were paid for that time, but it went toward answering the same generic prompt every session. Effort spent, but with nothing specific enough for a supervisor or researcher to act on in return.

EXPLORATION

Shortening the form or restyling it visually doesn't automatically make it feel less tedious. What really mattered was making the form more engaging and intuitive to actually answer.

RESTYLING THE FORM

My initial assumption was that making the form shorter, or at least keeping it the same length, was the way to re-engage tutors. Early explorations focused on making it feel shorter by design: reorganizing layout, cutting text density, and softening the visual monotony of the page.

Could questions adapt to the rating?

3 or fewer stars prompted improvement-focused questions, and 4+ triggered ones about what worked. This validated that progressive disclosure and dynamic, input-based questions both worked.

Could all students fit on one page?

I assumed that skipping page-by-page navigation would make the form feel shorter. But the questions were still generic, and seeing every student on one page felt more overwhelming.

What about tabs instead of scroll?

Splitting students across tabs was still the wrong structural fix. But visually, the chips softened the design, confirming that visual treatment still mattered even once the structural problem was solved elsewhere.

RETHINKING THE QUESTIONS

Reframing the problem around engagement shifted my focus from the layout to the questions. AI-generated contextual questions became the new default in place of free text, and I explored supporting changes around them, like option chips and a gentler thumb rating.

How can we map to the goal framework?

Effort, progress, and engagement each got their own option chips, followed by an AI-generated question grounded in those selections. This mapped to PLUS' goal-setting framework, but the range of options varied too much to fit a shared set.

What if we didn't use a rating at all?

In its place, the same three questions are answered through chips framed as positive, neutral, or negative, followed by a grounded AI question. This mapped to the framework and consolidated the range into a set that produced clean, trackable data.

What if tutors pick their own categories?

This version lets tutors choose what to write about with prompts appearing based on what they picked. This further validated the AI question direction: even with topic freedom, the underlying questions still felt generic.

SOLUTION

The final form is a structured, AI-assisted reflection tool that's easier to engage with, encourages more thoughtful responses, and gives supervisors and researchers better data in return.

Every piece works toward the same goal: context-grounded AI questions instead of generic ones, reflection mapped to PLUS' goal-setting framework, a clear path for real concerns, and a lighter, less frequent ask for self-reflection.

FEATURES

Context-grounded AI follow-up questions

The old session reflection paired text-heavy multiple choice with a free-response question that was identical every time. The redesigned flow replaces that with a single AI-generated follow-up question per section, grounded in the specific ratings and chip selections a tutor just made, so what the form asks is based only on that specific session.

Mapping to internal goal-setting framework

The student reflection now runs on PLUS' goal-setting categories: progress, effort, and engagement. Each is answered through chip options, followed by one AI-generated question grounded in those selections. Removing the rating cleared up ambiguity about what a tutor was even rating a student on.

Clarity in the escalation path

The old pathway was just a yes/no toggle without context. The redesign narrows escalation to what warrants immediate attention: mental health concerns, behavioral issues, and conduct concerns involving another tutor, for example. A "yes" or "other" selection pairs with a required description, and the options are shown as chips that surface examples upfront.

A Lighter, More Honest Self-Reflection

Previously required after every session, it now surfaces every 10th, easing the ask of tutors reflecting on themselves as often as they reflect on their students. The format changed too: a Likert scale gauging teaching confidence replaced the star rating, and the free response narrowed from open-ended to one positive and one area for improvement.

HOW IT WORKS

Tutor selections go in, rules pick the question, and a quality check decides whether it's asked at all.

Each AI follow-up runs the same underlying system, called separately for every section a tutor reaches: once for Session, once per selected Student, and once for Self on the sessions where it appears.

IMPACT

This redesign is approved and is currently moving through development. Check back for updates on real completion and data-quality metrics! ☺️