Yesterday, I was lucky to visit the Creative Computing Institute (CCI), and specifically Tim Smith‘s team, at the University of the Arts London. Tim invited me to give a short presentation about the research I am conducting at Royal Holloway under the supervision of Adam Ganz and Szonya Durant. Besides that, Tim showed me their facilities, and we had a talk about computer vision in cinema research.
In my talk, I presented a selection of the most interesting findings from the last 18 months: a) the difference between watching videos in vertical and horizontal frames, b) a null result regarding the effect of screen time on gaze patterns (both soon to be published in Projections), c) how people simultaneously watch narrative stimuli on TV and TikTok on a tablet (article just before submission), and d) which formal elements of narrative stimuli pull visual attention from a tablet to a TV screen (what I am working on). My main argument was that, contrary to my initial assumption, TikTok (and short video platforms in general) is probably not changing how we watch films. Rather, short videos (among other things) have made us aware of the audience’s inattentiveness. However, the fact that people are distracted is not new and it was always like that. It is just more visible now.
On a personal level, it was amazing to see the Creative Computing Institute and how Tim approaches research in the arts and humanities. I know very well from CoSTAR that arts research can look very different from what I was trained in. The CCI combines research activities with teaching, and it is inherently interdisciplinary. Truly, one can hardly say where one discipline ends and another begins. An illustrative example was when Tim showed me how they connected a virtual production studio with fNIRS and used it in a dance performance. The dancer’s brain activity produced a visualization on the LED panels behind him. This single example clearly shows how art can merge with neuroscience, technology, and programming, proving that these disciplines are not only communicating but also mutually beneficial.
During November, I visited several colleagues at various universities in England and while travelling by train, I read Adam Nixon’s book. Prestigious British universities and no-budget filmmaking – you may wonder why I would connect the two, but I will show you that there is a connection.
Nixon’s book is one filmmaker’s reflection on how filmmaking has changed and a report on the possibilities of film with virtually no budget. His perspective fascinates me. He is able to think about audiovisual culture in a very broad sense. In the preface, he prepares the reader for the fact that the best film of the present day may not be the work of a professional crew, let alone a film studio. We may already be living in an era where the best films are made on TikTok.
“Imagine an unknown artist is producing the best film of the current era right now. They are creating it outside centers of power, using a cheap camera with little expense. Cineastes are waiting to discover it. The question is, where will they find it? Will it debut on TikTok, Instagram, or at one of the thousands of fringe film festivals that dot the globe? The anticipation of this discovery keeps the cinema world exciting and ever evolving.” (p. viii)
Nixon’s experience as a film school teacher shows how the situation and function of schools in my field have changed. Whereas in the past they were places where students could access the best filming technology, today everyone has a 4K camera in their pocket. Access to technology alone does not make it easier to make a film. (Let’s leave sound aside.) This attitude is very similar to what my colleague from Olomouc, Tomáš Jirsa, says—students don’t go to school because of the availability of technology and don’t expect us to teach them how to use it. They have other reasons.
Nixon’s perspective shed a completely different light on what I saw at the workplaces I visited. Paolo Russo at Oxford Brookes University showed me around the campus, showed me the facilities, and we briefly discussed the technology available to their students. With Murray Smith at the University of Kent, we discussed his recent experience with interdisciplinary research and its applicability in teaching. Ceylin Ertekin from the London School of Economics showed me around their eye-tracking lab. Deborah Klika from the University of Greenwich gave me a tour of their studios, including virtual production. And everywhere I went, I thought to myself that it would take us a very long time to catch up in Czech film and humanities departments in general.
Nixon’s book puts this into a different perspective. Thanks to smartphones, we all have a camera in our pocket (in new iPhones, even with an eye tracker). Thanks to TikTok and YouTube, we all have access to a global audience. The democratization of technology may have completely redrawn the field in which we operate. The best Hollywood producers are suddenly meeting amateurs, and gatekeepers are losing power.
Don’t get me wrong. I still think technology is essential. The fact that Royal Holloway and the University of Greenwich have virtual production studios means that they have enormous opportunities in terms of research. The excellent sound technology at Oxford Brookes University offers unparalleled opportunities for their students, even if they use it to shoot on their phones. But if we stick to the educational dimension and student expectations, Nixon is probably right and access to the latest technology is not essential. After all, the best films are already being made for TikTok.
For me, Nixon’s book is mainly an inspiration to rethink what film studies are for and what their place is in today’s education system.
Today and tomorrow, I am attending the ZIP-SCENE Conference in Prague. The conference is part of the Art*VR Festival at the DOX Center for Contemporary Art.
When I saw the program, I was surprised that my proposal was accepted at all. Most of the presentations are about VR, XR, video games, etc. So my 2D vertical videos stand out a bit. I tried to choose a framing that placed the viewing of vertical videos in the context of the audiovisual landscape, which also includes VR. Some colleagues certainly noticed this, but they were kind enough to keep their doubts about compatibility to themselves.
On the contrary, I received positive feedback, and the main finding that vertical framing causes fragmented viewing surprised several colleagues. (I would like to write more here, but only after we publish it in an article we are preparing with Szonya Durant and Adam Ganz).
Nevertheless, from the other contributions, I had the impression that I had found a similarly minded audience. The ZIP-SCENE Conference is a great achievement by Ágnes Karolina Bakk‘s team. This is already the seventh year, and they have succeeded in bringing together artists and researchers from various fields who think not only about what they create, but also about how it affects the audience.
I was most interested in the following presentations:
Niels Erik Raursø presented the research activities of the Augmented Performance Lab at Aalborg University in Denmark. They are investigating how to use EEG to analyze audience responses to storytelling. In theory, this could lead to greater personalization of storytelling for the needs of individual viewers in the future. I would probably need to hear more about their research, because in this 15-minute presentation I identified some methodological problems. For example, I think they could easily fall into the trap of choosing the wrong way to segment the narrative. In any case, I found Niels’ research extremely inspiring.
Another thought-provoking presentation was by Felix Carter, Iain Glichrist, and Danae Stanton-Frasert. They experimentally explored how a change in the narrative affects the audience’s attention towards exogenous cues. In my opinion, their research potentially shows how big a difference there can be in how we understand films. Not that we encounter different editing of films in cinemas (although that is also possible). I am referring to the simple fact that not all of us pay full attention to films throughout their duration. We look at our mobile phones while watching a film, we go to the fridge or the toilet, we talk to people around us, and films on TV are interrupted by commercials.
The last contribution I want to mention from today’s program was presented by Pavel Srp. He talked about an experiment in which they tested whether a linear or logarithmic function is more suitable for determining the distance of a sound source in relation to volume. The conclusion was that a logarithmic function is more suitable. Pavel’s paper is the third I have heard at the conference in a short time from “sound people,” and all three were excellent and provided great insight into audience perception.
And since I was at the Art*VR festival, I wanted to try out what it’s like to be in virtual reality. In short, after five minutes I felt sick and it took me about an hour to recover. But tomorrow I’ll give VR another chance.
In recent weeks, the boom in generative AI models has begun to shift from text and images to video. Anyone can create an eight-second video using Google’s Veo model. Open AI has announced the launch (currently only in North America and by invitation) of a video platform running on the Sora 2 model. In addition, other platforms are emerging that enable the generation of special forms of content, such as micro dramas.
In our family, video generation has become part of our daily bedtime routine. Every evening, the children simply dictate to me what they want their fairy tale to be about. Gemini first prepares the text for me, and after I finish reading it, we generate the video. I consider this to be one of the greatest advantages of generative AI, as the quality of the fairy tales does not depend on how tired I am.
The problem, of course, is that eight-second videos created in two minutes enchant us now, but in a moment we will consider them the norm and they will no longer be enough for us. And if I understand correctly, the problem for developers is still to keep the characters looking consistent between shots, or to ensure that the resulting videos follow basic cinematic continuity rules. This is not surprising, because AI must first learn to tell stories like filmmakers in order to match them, and it took filmmakers decades to do so. And developers could be helped by a traditional institution well known to film historians – the archive – with its tried and tested rules of film storytelling.
I realized this in connection with The Development of Generative Artificial Intelligence from a Copyright Perspective from May 2025, which was brought to my attention by my friend and lawyer Miroslav Obernauer. The creation of a license for training generative AI is being considered. This would solve the problem of the need to train AI on existing works on the one hand and copyright protection on the other.
And this is where there is room for film archives to become centers for training AI, in addition to their role of preservation, restoration, and dissemination.
Take, for example, the Národní filmový archiv in Prague. Archivists care for hundreds of films made in Czechoslovakia – films that could be used to train generative AI. And their collections include thousands more films from other countries around the world.
Of course, I don’t know if this will happen. But I like the idea that an institution such as an archive, which – let’s be honest – is not one of the driving forces behind the digital revolution and does not exactly come across as cool to the public, could find a new role and become integral part of the creation of modern technologies.
After six months of work on this project, I finally have my first set of data. In this post, I want to reflect on what these data might tell us about people’s relationship with TikTok—and what they reveal about how we study the perception of this platform. It’s not just about what the data show, but also how we interpret them and what questions we should be asking next.
In my initial experiment, I collected valid data from 28 participants. Each participant watched the same set of videos on a computer with eye-tracker, presented in two formats: landscape and portrait. Before every session I made a note of participants’ screen time (TikTok, Instagram, and YouTube) – not through self-reporting, but by directly checking the screen time data on their phones.
After processing and analysing the data, we obtained interesting information about differences in fixation entropy, saccade directions, and, in general, the relationship between video format and stylistic elements in the context of perception. We are preparing academic outputs on all of this. But what surprised me the most was something else.
We were unable to measure any convincing effect of screen time on the way videos are watched. It even makes me rethink some of my hypotheses. But let’s take it step by step.
To be honest, I did find one statistically significant correlation. And it’s one that deserves a headline in the newspapers: TikTok makes you less focused when watching movies!
Sounds serious, right?
More precisely, we found that participants with higher average screen time on TikTok have a greater distance between fixations when watching landscape videos. In other words, if you watch a lot of TikTok and then watch a movie, your eyes will jump around more than the eyes of people who don’t watch TikTok.
Does it still sound serious?
Take a look at the graph illustrating this effect. The p-value even shows that the effect is statistically significant.
Does it sound even more serious now?
Notice the dots on the y-axis that have zero screen time on TikTok. Let’s try visualizing the data differently to get two groups: TikTok users and non-users.
And suddenly, the effect is gone. Or at least it is so small and statistically insignificant that nothing can be said with certainty on its basis.
What happened?
Imagine that you collect a huge amount of data (fixations, saccades, pupil diameters, blinking… in the spreadsheet with more than 100 columns and calculate metrics such as entropy, dispersion…) and add to it other data from a questionnaire (age, gender, average screen time…). Then you just try to look for what correlates with what until you find that some correlation is statistically significant. In other words, you are committing fraud by randomly comparing data and looking for something that looks like positive result. And when you get a low p-value, only then do you formulate a hypothesis. This is called data fishing or p-hacking.
I was in the opposite situation. I measured the effect predicted by the hypothesis. The problem was that when I then explored other variants of correlations between screen time and eye-tracker data, I didn’t find any other statistically significant correlations. The one with screen time on TikTok and fixation distance in landscape videos was the only one. Sometimes it even showed me that higher screen time correlates with more focused fixations. The complete opposite.
These findings have led me to reconsider some of my initial assumptions. I originally believed that long-term exposure to TikTok would affect the way we watch films. But based on the current data, that effect is either not present or not captured by my experimental design. In any case, I don’t have sufficient evidence to claim that screen time directly influences how we watch videos. What’s more, I began to doubt whether it still makes sense to continue focusing on studying influence of screen time on viewing habits.
Another important takeaway concerns how we interpret claims about TikTok’s influence on cognition. Keeping in mind what Stuart Ritchie wrote in his excellent book Science Fictions, I’ve begun to question the many articles that describe TikTok as fundamentally reshaping our brains and behavior. Perhaps those effects are real—but we should be cautious when such strong claims are supported by just a single metric from an eye tracker or similarly narrow data sources.
Finally, these results also shape how I’ll present my own research moving forward. I’ll keep the question in the title of this post, because it’s accessible and engaging. But in answering it, I’ll emphasise a crucial distinction between a) TikTok as an app, b) the formal elements typical of TikTok videos, and c) the vertical video format itself. Unlike screen time, both the formal elements and the vertical format show clear evidence of influencing how we watch