Filler sounds and words are commonly used as they facilitate breaking up speech. It allows people to think, take a break and process information. Anecdotal evidence suggests some university students could feel the quality of teaching is negatively affected by excessive filler use. This research aims to address what filler words are used, their frequency and the best available methods to accurately assess filler use.
Nine participants from the School of Biomedical Sciences (SoBS) and the Plymouth Business School (PBS) each provided two Panopto lecture recordings, one was perceived to be delivered well, and the other could have been delivered better (e.g. not main subject area or lack of preparation time). Ethical approval was given by the University of Plymouth Faculty of Science and Engineering Research Ethics and Integrity Committee. All participants completed a consent form using the JISC online survey platform which was shared via a link or QR code. We analysed six of the participants recordings manually, counting the number of various filler words (um, so, okay, uh, like etc.) and any prominent patterns. All participant data was also analysed using machine learning (ML). By creating a labelled dataset of the filler words including positional context of words, we were able to use DistilBERT, a smaller and faster version of BERT (Bidirectional Encoder Representations from Transformers), which is a machine learning model made to understand human language. We achieved an F1 score (combined precision and recall) of 0.97 for predicting filler words in text and were able to quickly and accurately count the number of filler words, using the captions generated on Panopto. Mean (SD) and paired t-testing were used for data analysis. P values of <0.05 were considered statistically significant.
Turning to machine learning aided to reduce the manual time in recording fillers, as human assessment took approximately 1.5 to 2.5 hours for a one-hour lecture compared to 20s for the ML solution (with 2 minutes of training on a laptop GPU). Mean (SD) manual filler count (n=6) was 892.8 (389.6) compared to a mean ML filler count (n=6) of 1006.1 (430.6). The mean manual filler ratio (percentage of filler count with respect to total words) was significantly lower than the ML equivalent (6.6 vs. 7.4%; p=0.008). Additionally, we found the ML filler ratio to be significantly higher for the poorly delivered lectures, contrasted to the well delivered ones (n=9; p=0.005). However, the filler count per minute for both lectures showed no significant difference (p=0.38).
Our findings show that the ML solution was incredibly efficient and was approximately within 10% of manual data recordings. We also noted a potential difference between lecture quality and filler use; however, these findings require further research into the factors affecting filler use. We hope the results achieved will provide more insight into filler use during teaching, in order to better inform educators in HEIs. These findings may also increase awareness of filler overuse for teaching staff.