Evidence hub · Innovation · Impact · Depth

Educational outcomes made possible

The intended outcome is not faster answers but more capable participants. Students progress from multimodal chat, to tool-using AI, to supervised agents while learning to explain and verify outputs. Teachers compare models, create materials and test course questions. Teaching assistants use shared guidance and, in the early-stage GradePilot workflow, review rubric-guided drafts while retaining authority to edit, reject and decide every grade.

Diagram of role-specific AI literacy for students, teachers and teaching assistants, with verification and human judgement at the centre.
Figure 1. One ecosystem, role-specific outcomes. Students verify, teachers design and teaching assistants review; human responsibility remains common to every role.

These outcomes are difficult to reproduce through occasional office hours or a generic chatbot. They depend on repeated access, role-specific learning, realistic demonstrations, inspectable tool use and evidence from actual courses. Shared credentials and free access let participants practise regardless of their ability to buy a commercial subscription.

Achieved engagement

Direct platform records show:

  • More than 300 active users across at least ten Physics and Statistics courses
  • Approximately 100 daily active users
  • More than 60 billion tokens served
  • More than 100 written feedback responses

Counts from different activities may overlap and are not summed as unique beneficiaries. The customised OpenWebUI meters each student’s expenditure and applies a weekly cap, supporting repeated use without shifting subscription costs to learners.

Impact summary showing more than 300 active users, approximately 100 daily users, more than 60 billion tokens, more than 100 feedback responses, at least ten courses and weekly student budget caps.
Figure 2. Sustained use, not registration alone. Platform records combine reach, return use, service volume, course coverage and cost controls.

The work also engages CUHK staff, professors at other universities and secondary-school teachers. Planned access across the Faculty of Science is expected to bring the connected services to more than 1,000 users. This is a projection, clearly separated from the achieved figures above. GradePilot is a working early-stage prototype and is also excluded from achieved-impact totals.

Evidence of learning and usability

The 2024–25 evaluation cycle collected more than 100 written responses. More than 75% of respondents rated the service easy to use, and more than 80% reported improved learning. These are reported survey outcomes, not a claim that every new AI capability automatically improves education.

Working examples demonstrate the learning surfaces behind those results. The tool-using environment exposes sources and executed calculations, so a learner can inspect the reasoning path. The agent environment exposes the prompt, task boundary and response. GradePilot places the source submission, rubric fields, draft-assistance control and final save action on one review surface.

Authenticated OpenWebUI response showing official-source web research, Code Interpreter execution, equations and a results table.
Figure 3. Verifiable tool use. Official sources, computation and results can be checked within the same educational interaction.
Authenticated GradePilot teaching-assistant grading screen showing score fields, feedback, source and PDF panes, AI Fill and Save controls.
Figure 4. Feedback support with retained TA authority. GradePilot is evaluated as an early-stage human-reviewed workflow; it does not make or submit the official grading decision.

Monitoring and ongoing improvement

Usage dashboards track adoption, return use, activities and per-student cost. Surveys and open feedback examine ease of use, learning benefit, course relevance and over-reliance. Each service is also tested on representative educational tasks: whether students can explain and verify outputs, teachers can select suitable tools and design appropriate activities, and TAs can identify errors, supervise agent work and review drafts without surrendering judgement.

Continuous improvement loop connecting usage dashboards, surveys, representative tasks, safeguards and revisions.
Figure 5. Evidence leads to revision. Usage and learning evidence inform changes to services, tutorials and classroom guidance for the next cohort.

Future GradePilot evaluation will compare its provisional drafts with TA decisions and document edits or rejections. Consent, anonymisation, restricted access and retention are reviewed alongside educational effects. Repeated measures across cohorts will test growth in capability rather than traffic alone.

Evidence hub · Innovation · Impact · Depth