Category Archives: Uncategorized

Dear Data,…

Valentine’s day week for the international data community is not only a time for expressing your love to the significant others in your life. As it is also Love Data Week, it is also a time to reflect on your love for all things data! That was the goal for the Research Data team this year! The theme of this year’s Love Data Week was “My Kind of Data”, suggesting that data workers – researchers and analysts alike – have a relationship to data that is personal, often idiosyncratic, and almost always heartfelt. The Research Data team, as supporters of the University’s researchers, are interested in such relationships and are always eager to discover the distinctive needs that the disciplinary differences between the University’s departments create. This year, the Research Data team decided that they wanted to find out from students and researchers from the Arts, Humanities and Social Sciences (AHSS) what was their kind of data.

To do so, the Research Data team positioned themselves at the Foyer of the Alison Richard Building on the University’s Sidgwick Site, which is home to several AHSS departments, for two mornings on Monday the 12th and Thursday the 15th of February. Across the city, Data Champion Lizzie Sparrow was leading the charge with science, technology, engineering, mathematics (STEMM) students and researchers by holding her own pop-up at the West Hub. Like the Research Data team, and as a Research Support Librarian (Engineering) herself, Lizzie is also interested in the relationships that researchers have with data. Her approach, however, would likely be different. Unlike researchers in the STEMM subjects, the term data for AHSS students and researchers can sometimes feel exclusionary as they may not consider what they generate through research as data. From our perspective on the other hand, any material that goes on to form any part of their research is one’s data. To bring attention to this, the team tried to engage passers-by with the provocation “you have research data, change our minds!” The provocation was successful and many conversations were had on the different ways that members of the Sidgwick community understood data in their research.

The Research Data Team from the Office of Scholarly Communication (Cambridge University Library), from left to right: Clair Castle, Lutfi Othman, Kim Clugston.

The team was pleased to find that there was a general interest in the services of the Research Data team among the Sidgwick community, and we were happy to be able to share with others how we can help them with their data management and planning.

Some treats for those who stop by.
Our Open Research poster, designed by Clair Castle.

The team tried to capture the sentiments of the conversations had by asking the Sidgwick community to partake in 2 short activities as they departed our pop-up to better understand  their relationship with data (in exchange for Love Hearts sweets!). Firstly, we asked them to describe to us what data was to them, a question that we are extremely fond of asking! As usual, the answers were informative and they helped us to gain a sense of the varying data types that the Sidgwick community worked with – from political tracts and archival materials to balance sheets and land deeds from the early modern era.

Activity 1: Lots of different data types in the AHSS community!

For the second activity, we asked them what term best captured the materials that formed the basis of their scholarly work: data, research materials, or other? To our surprise, the majority of people we spoke to over both days saw themselves as working with data, more than double the number that saw themselves working with research materials, with a small number seeing themselves as working with both, interchangeably. This finding illustrated something that has been increasingly discussed in the Research Data team office: that finding alternatives to the term data may make our services and initiatives more appealing to members of the AHSS community. This is something we will take into account when targeting our outreach in the future. Yet, one thing is certain – our Research Data services are needed by the AHSS community just as much as it is by the STEMM community.

Activity 2: More generators of ‘data’ than we expected!

The pop-ups at the Alison Richard building were encouraging and it is hoped that fruitful relationships will transpire from these events. This is something that we may hold again soon. It was a good way to communicate our message and make others aware of the services of the Research Data team. Over at the West Hub Lizzie was not as encouraged, having only managed to have in depth chats with a couple of people. She reported that lots of people were very determinedly on their way somewhere and not up for stopping to talk. The time and/or location did not seem right for the intended audience. I suppose, we shouldn’t stand in between a student and their food. In any case, there were lots to take away from this Love Data Week pop-ups, and lots to reflect when we plan for our next pop-up, be it for Love Data Week 2025 or just as a periodic service to the research community here at Cambridge. Perhaps when the weather is nicer in the summer, we will do a pop-up outdoors in the middle of the Sidgwick site, or at research events throughout the University. If you have any ideas on where it would be good for us to hold such a pop-up, do let us know!

Mapping the world through data – The November 2023 Data Champion Forum 

The November Data Champion forum was a geography/geospatial data themed edition of the bi-monthly gathering, this time hosted by the Physiology department. As usual, the Data Champions in attendance were treated to two presentations. Up first was Martin Lucas-Smith from the Department of Geography who introduced the audience to the OpenStreetMap (OSM) project, a global community mapping project using crowdsourcing. Just as Wikipedia is for textual information, OSM results in a worldwide map created by everyday people who map the world themselves. The resulting maps can vary in terms of its focus such as the transport map, which is a map which shows public transport lanes like railways, buses and trams worldwide, and the humanitarian map, which is an initiative dedicated to humanitarian action through open mapping. Martin is personally involved in a project called CycleStreets which, as the name implies, uses open mapping of bicycle infrastructure. The Department of Geography uses OSM as a background for its Cambridge Air Photos websites. Projects like these, Martin highlighted, demonstrate how community gets generated around open data. 

CycleStreets: Martin at the November 2023 Data Champion Forum

In his presentation, Martin explained the mechanics of OSM such as its data structure, how the maps are edited, and how data can be used in systems like routing engines. Editing the maps and the decision-making processes that go behind how a path is represented visually on the map is the point where the OSM community comes to action. While the data in OSM consists primarily of geometric points (called ‘Nodes’) and lines (called ‘Ways’) coupled with tags which denotes metadata values, the norms about how to define this information can only come about by consensus from the OSM community. This is perhaps different to more formal database structures that might be employed within corporate efforts such as Google. Because of its widespread crowdsourced nature, OSM tends to be more detailed than other maps for less well-served communities such as people cycling or walking, and its metadata is richer, as they are created by people who are intimately familiar with the areas that they are mapping. A map by users for users. 

Next up was Dr Rachel Sippy, a Research Associate with the Department of Genetics who presented how geospatial data factored into epidemiological research. In her work, the questions of ‘who’, ‘when’, and ‘where’ a disease outbreak occurred are important, at it is the where that gives her research a geographical focus. Maps, however, are often not detailed enough to provide information about an outbreak of disease among a population or community as maps can only mark out the incident site, the place, whereas the spatial context of that place, which she denotes as space, is equally as important in understanding disease outbreaks.  

Of ‘Space’ and ‘Place’: Rachel at the November 2023 Data Champion forum

It can be difficult, however, to understand what a researcher is measuring and what types of data can be used to measure space and/or place. Spatial data, as Rachel pointed out, can be difficult to work with and the researcher has to decide if spatial data is a burden or fundamental to the understanding of a disease outbreak in a particular setting. Rachel discussed several aspects of spatial data which she has considered in her research such as visualisation techniques, data sources and methods of analysis. They all come with their own sets of challenges and researchers have to navigate them to decide how best to tell the fundamental story that answers the research question. This essentially comes down to an act of curation of spatial data, as Rachel pointed out, quoting Mark Monmoneir, that “not only is it easy to lie with maps, it’s essential”. In doing so, researchers working with spatial data would have to navigate the political and cultural hierarchies that are explicitly and implicitly inherent to places, and any ethical considerations relating to both the human and non-human (animal) inhabitants of those geographical locations. Ultimately, how data owners choose to model the spatial data will affect the analysis of the research, and with it, its utility for public health. 

After lunch, both Martin and Rachel sat together to hold a combined Q&A session and a discussion emerged around the topic of subjectivity. A question was raised to Rachel regarding mapping and subjectivity, as it was noticed that how she described place, which included socio-cultural meanings and personal preferences of the inhabitants of the place, can be considered to be subjective in manner. Rachel agreed and alluded back to her presentation, where she mentioned that these aspects of mapping can get fuzzy as researchers would have to deal with matters relating to identity, political affiliations and personal opinions, such as how safe an individual may feel in a particular place. Martin added that with the OSM project the data must be objective as possible, yet the maps themselves are subjective views of objective data.  

Rachel and Martin answering questions from the Data Champions at the November 2023 forum

Martin also brought to attention that maps are contested spaces because spaces can be political in nature. Rachel added that sometimes, maps do not appropriately represent the contested nature of her field sites, which she only learned through time on the field. In this way, context is very important for “real mapping”. As an example, Martin discussed his “UK collision data” map, created outside the University, which states where collisions have happened, giving the example of one of central Cambridge’s busiest streets, Mill Road: without contextual information such as what time these collisions occurred, what vehicles were involved, and the environmental conditions at the time of the accident, a collision map may not be that valuable. To this end, it was asked whether ethnographic research could provide useful data in the act of mapping and the speakers agreed. 

The Data Picture

I was recently named one of “the next generation of [library] leaders” as part of the CILIP 125, having been recognised as an individual who contributes energy and knowledge to improving and impacting their organisation. My area of expertise, and thus recognition, lies with the use of data within libraries. As a data analyst for the Office of Scholarly Communications at Cambridge University Library, my role focuses on empowering decisions with data driven understanding – such as supporting the Springer Nature negotiations. To develop my understanding of data, and its role within a wider organisation, further, I engage with data beyond the library – such as the Big Data London conference and the Carruthers and Jackson Data Leaders’ Summer School. Reflecting on the use of data in the wider world, what can be expected of the library and data?

The summer school provided practical advice, proven methodologies, and guidance that could apply across a variety of businesses. The course is designed to provide insight on the workflow of data officers, and their role within an organisation – no matter its stage of data maturity and literacy. Over the course of the ten weeks, leading experts discussed the role of a chief data officer (CDO), both as a business development opportunity, and as a career path for individuals. It explored the risk and governance of data within an organisation, and the final weeks focused strongly on the role of people and teams associated with data.

Peter Jackson and Caroline Carruthers addressed the differing types of CDO and described a pendulum between ‘risk aversion’ and ‘value added’. Understanding the balance between secure and proper data governance (GDPR for example) and providing value through data (such as setting up automation). The pendulum of risk to reward is relevant to many roles, including those within the library. Understanding the need to divide time and energy between creating policies and getting decision making results, is just as relevant to my role as a chief data officer. In my role I have supported decision making staff through data production, but equally, to instil a culture of data, time and energy must be dedicated to risk aversion, through tasks of researching data management, preparing training sessions for data storage, and supporting staff in data preparation.

Another important concept introduced was the DIKW pyramid – Data, Information, Knowledge, Wisdom – for understanding the value created from data. The base of the pyramid is (raw) Data, which can be processed into (useful) Information. This Information is data with meaning and a purpose and can be organised into (insightful) knowledge. Knowledge combines experiences, values, insights, and contextual information, which can then transcend to (integral) Wisdom. Wisdom is considered a deeper understanding with ethical implications and the ability to define ‘why’. The DIKW pyramid provided a frame of thought for presenting and approaching future data projects. Understanding the requirement to provide, data, information or knowledge, to better support a decision-making team.

To develop communication skills, expert Scott Taylor, known as The Data Whisperer, spoke about the three V’s for data storytelling: Vocabulary, Voice and Vision. Combining an accessible vocabulary, with a common voice will illuminate the business vision, and why that is important. This overarching concept for an organisations data approach can be scaled down to support individual data workers, to provide value – which should either grow, improve or protect the business case. Understanding how to communicate the data is a key skill as “Hardware comes and goes, software comes and goes, but data remains”. And that data that remains should be used to either grow, improve or protect the business, such that data gathered should be usable data!

At Big Data London, the organisation Women in Data hosted conversations about nurturing a culture of learning within data teams. Pulling from their experiences from minority backgrounds, the speakers highlighted the power in upskilling, sharing skills across teams and being an advocate on oneself and skills. As for what to upskill, data literacy was a hot topic across the conference. Data literacy, also called data fluency and data confidence, is the combination of ability, skills and confidence surround data and its uses. Data literacy enables more efficient work, and begs the question, what is the base level of data literacy / confidence across the library? Librarians use data daily; checking in/out material, answering students’ queries, or tracking the use of space, but are all librarians confident to use that data? This is an area I hope to explore further at the CUL, to ensure staff can use the data they have to support decisions.

Engaging with the world of data provides a big picture of the possibilities within the library. Conversations of AI (Artificial Intelligence), data policies and maturity, and shiny-new databases, software, and services, demonstrate the growing adoption of data, and therefore, libraries should follow suit. Actively taking snippets of larger conversations, developing ideas within the library space, and exploring the possibilities with data will help libraries thrive in this world of technological growth.