Clearing is open! Find out more or get in touch 01604 214808.

Fast answers, slower judgement: Using Copilot to find evidence in university reports

Date 4 September 2026

Nicola Denning writes about how she used Microsoft Copilot to identify useful patterns across a large collection of reports and to suggest where to look next.

Nicola Denning

Digital Horizons Hub logo. Shows 4 different colours circles connected to lines. Looks like a switchboard.Generative AI is often discussed in relation to teaching and assessment, but it can also support the less visible work that shapes a university curriculum.

Nicola Denning works in learning design at the University of Northampton. When colleagues needed evidence about how mental health and positive wellbeing were being considered in curriculum development, she wondered whether Microsoft Copilot could help.

The information was not held in a ready-made dataset. It was scattered across a large collection of reports from course validation events. Reading every report manually would have taken a considerable amount of time, so Nicola decided to see what Copilot could find.

Her experience showed how quickly AI can identify useful patterns. It also revealed why speed must be accompanied by careful human judgement.

Starting with information the University already held

Validation reports record the discussions and decisions that take place when courses are approved or reviewed. They may include examples of effective curriculum design, as well as recommendations where a panel believes further work is needed.

Nicola thought these reports might contain relevant evidence about wellbeing, even though they had not been written for that purpose.

She gathered the confirmed reports into a folder and directed Copilot to examine them. Her prompt asked it to identify references to mental health and wellbeing, with a particular focus on positive practice.

“There were a lot of reports, so Copilot seemed like a good place to start. I am not saying it was the be-all and end-all, but it gave me a way into the material.”

Within a short time, Copilot suggested themes and linked them to examples from courses. This gave Nicola an initial overview without requiring her to read every report from beginning to end.

The process was quick to set up because she already had authorised access to the documents and Copilot was available within the University’s digital environment. There was no need to transfer the reports to an external system.

The moment Copilot got it wrong

The results initially appeared plausible. Some referred to validation events with which Nicola was already familiar, so she could see that Copilot was working within the right general area.

She then followed the links back to the original reports to check the evidence for herself. Most of the examples she sampled were broadly accurate. One was not.

“In one case, it simply was not what the report was saying. It was almost the complete opposite, which was interesting, and alarming.”

This changed the meaning of the exercise. Copilot could not be treated as a reliable analyst whose conclusions were ready to pass on. It was better understood as a tool for locating possible themes and suggesting where a person should look next.

The inaccurate example also demonstrated how difficult AI errors can be to spot. The output did not necessarily look obviously false. It became apparent only when Nicola returned to the source and read it in context.

“The main challenge with generative AI is whether what you are getting is reliable.”

Useful for finding a starting point

Despite that error, Nicola found the exercise worthwhile. Manually searching the full set of reports would have taken far longer, particularly because she was looking for broad patterns rather than one known phrase.

Copilot provided a rapid first pass through the material. It also suggested possible lines of inquiry that Nicola might not have considered at the outset.

This capacity to surface an unexpected connection can be useful, but it also needs to be treated cautiously. Generative AI is designed to produce a helpful response, which can lead it to make the available material fit the question too neatly.

As Nicola put it: “AI aims to please. We have to be careful that it is not sending us something too rogue because it thinks that is what we want to hear.”

The value therefore came from combining machine speed with human scrutiny. Copilot reduced the amount of material that needed to be examined initially. Nicola then applied her knowledge of learning design to decide which findings were meaningful and which needed further checking.

The real skill is not opening the tool

Using Copilot required very little technical expertise. The more demanding skills came before and after the prompt.

The initial question had to be clear enough to guide the search. Nicola also needed to specify what kind of output would be useful. Without that detail, Copilot might provide a response that required substantial reworking.

“Anybody can put in a question. The skill comes in the delicacy of the prompt and in the interpretation of what you get back.”

If repeating the exercise, Nicola would spend more time refining the prompt. She would define the boundaries of the task more precisely and tell Copilot how to organise its response.

Yet even an excellent prompt would not remove the need for checking. The user still needs to ask whether the findings make sense and whether they can withstand scrutiny.

This is especially important when AI-supported analysis may contribute to institutional decision-making. A convincing summary can influence how a problem is understood, even when its interpretation of the evidence is weak.

Keeping human decisions at the centre

The reports did not contain personal identifying information, so Nicola did not encounter concerns about individual privacy in this case. Her main ethical concern was the possibility of passing inaccurate information into work connected with University strategy.

“The human decisions need to be there. People have to decide whether something looks accurate and whether it is something we should integrate into our work.”

Nicola is not particularly fond of the phrase “human in the loop”, but the principle fits her experience. Copilot gathered and organised information, while responsibility for interpreting it remained with the person using the tool.

This was not simply a final check added at the end. Human judgement shaped the question and guided the investigation. It also determined whether the output was sufficiently trustworthy to inform later work.

Modelling careful AI use

For Nicola, the lessons extend beyond learning design. Teaching staff can help students understand generative AI by using it openly and discussing what happens.

This means showing where AI saves time while also making its limitations visible. A confident answer should not be confused with a reliable one.

“It is important for students to see AI use modelled and to hear about the difficulties as well as the benefits.”

Students are likely to encounter AI in their future workplaces, although its role will vary between professions. Universities therefore need to help them develop more than the ability to generate an answer. They need opportunities to question outputs and decide how far those outputs can be trusted.

Nicola describes herself as someone who simply had a go, rather than an AI expert. That makes this example especially useful. Responsible experimentation does not require complete technical mastery. It begins with a real problem and a willingness to test what the tool can offer.

The result may save a great deal of time. It may also be wrong. The important part is knowing that both possibilities can be true at once.

Nicola Denning
Nicola Denning

Nicola Denning works in learning design at the University of Northampton.

Subscribe to get the latest about our projects