Annalena Alber
Co-Creation of Visual and Textual Meaning. Changes in Image-Text Relations Across Centuries of Travel and Mountaineering Literature
Teaser:
Travel writing has always been a duet of image and text — each shaping how the other is read. But how exactly do the two modalities share meaning, and how do those relationships shift across languages, audiences, and centuries of technological change? Drawing on Barthes' anchorage and relay and Panofsky's situated viewer, this project turns to an unexpected ally: vision-language models. Rerunnable under controlled conditions, yet steeped in cultural and technological bias, these models become both instrument and object of analysis — their failures on historical material exposing weaknesses in themselves.
The study spans two remarkable corpora: 450 years of Italian travel literature at the Bibliotheca Hertziana and the alpine mountaineering texts of the Text+Berg collection at the University of Zurich. From narrated journeys and drawn panoramas to competing vedute and itineraries, the goal is a structured framework for analysing how image and text co-create meaning in historical corpora.
Project Descrption:
In much of travel writing, images and texts have been created in tandem, the reading of one influenced by the reading of the other. How this interrelationship varies based on the involved language communities, the represented sceneries, the changes in creators and consumers, the technological advances surrounding the media remains difficult to analyse. The problem with jointly analysing the modalities mainly lies in their underdetermination, since neither can be isolated as carrying the complete meaning on its own.
This difficulty raises the question of what an image can say that a text alone cannot, what a text can say that a single image cannot. It raises the question of how an image shapes a text and how a text reframes an image, and how the two modalities can take on each other’s functions and influence each other’s movement. It raises the question of how the relations change, depending on the context but also the diachronic evolution of text and image creation itself.
The variability of this relationship can be addressed using Barthes, who in “Rhetoric of the Image” (1977) distinguishes two ways that the text can influence the reception of an image. Anchorage, on one hand, fixes the meaning of an image through the accompanying text, whereas relay describes a relationship where the meaning arises through the co-existence of the two. Panofsky in “Meaning in the Visual Arts” (1955) adds a constraint: no reading or interpretation can occur in a vacuum. A human reader brings their own evolving knowledge and historical situatedness to the reading, which cannot be ignored. A text cannot be un-read, an image cannot be un-seen.
This project aims to analyse vision-language models in such a context. Unlike human readers, the models can be rerun under controlled conditions, pre-existing cultural knowledge averaged into them; model and material can be crossed. The goal is to create dual descriptions, one linked to the text and one separated from it. Neither description is neutral but rather consistent in its biases, as the model is shaped by cultural and technological influences, making their differences a test for the model and exposing its weaknesses through large-scale systematic corpus-linguistic and vision-based evaluations. Through this, the model becomes both the instrument and the object of analysis. Where the model fails, particularly on historical material, it highlights issues within itself.
In this project, the focus lies on movement and space, written and visualized through drawn or narrated journeys, engraved or described panoramas, from plains and cities up to the mountains; this is where the modalities compete for content with each other, where veduta and itinerary, spatial narrative and cartography naturally meet. The analyses draw on two collections of articles, travel accounts, and recollections of journeys, spanning over 450 years: Italian travel literature, housed at the Bibliotheca Hertziana in Rome, and the Text+Berg alpine mountaineering corpus, housed at the University of Zurich. While the respective linguistic traditions, geographical locations, and publication practices differ, image and text focus on the same themes. In both corpora, image and text oftentimes delineate the same path, the same building, the same mountain, thus enabling the analysis of their cross-modal creation of meaning.
The overarching goal of the project is to develop a structured framework for analysing visual and textual meaning in historical corpora.
Bio
Annalena Alber is a Predoctoral Fellow with the IMPRS-MDH, starting September 2026. She has previously completed a master's in Digital Linguistics at the University of Zurich, where she focused on multimodality and accessibility. She completed her bachelor's degree at the University of Innsbruck in Linguistics and Digital Science. During her repeated work with collections and corpora she noticed an asymmetry: modalities were seldom treated equally in digital editions, thus creating structural gaps, mostly invisible unless the physical material was inspected. Her work aims to make material structurally accessible, both to researchers and to computational analyses, and to develop new methods for working with multimodal corpora. For her thesis she is working with multimodal travel and mountaineering literature, researching changes in image-text relationships across the evolution of imaging techniques, using vision-language models both as tools and as objects of analysis.