Category Archives: Uncategorized

Mapping the world through data – The November 2023 Data Champion Forum 

The November Data Champion forum was a geography/geospatial data themed edition of the bi-monthly gathering, this time hosted by the Physiology department. As usual, the Data Champions in attendance were treated to two presentations. Up first was Martin Lucas-Smith from the Department of Geography who introduced the audience to the OpenStreetMap (OSM) project, a global community mapping project using crowdsourcing. Just as Wikipedia is for textual information, OSM results in a worldwide map created by everyday people who map the world themselves. The resulting maps can vary in terms of its focus such as the transport map, which is a map which shows public transport lanes like railways, buses and trams worldwide, and the humanitarian map, which is an initiative dedicated to humanitarian action through open mapping. Martin is personally involved in a project called CycleStreets which, as the name implies, uses open mapping of bicycle infrastructure. The Department of Geography uses OSM as a background for its Cambridge Air Photos websites. Projects like these, Martin highlighted, demonstrate how community gets generated around open data. 

CycleStreets: Martin at the November 2023 Data Champion Forum

In his presentation, Martin explained the mechanics of OSM such as its data structure, how the maps are edited, and how data can be used in systems like routing engines. Editing the maps and the decision-making processes that go behind how a path is represented visually on the map is the point where the OSM community comes to action. While the data in OSM consists primarily of geometric points (called ‘Nodes’) and lines (called ‘Ways’) coupled with tags which denotes metadata values, the norms about how to define this information can only come about by consensus from the OSM community. This is perhaps different to more formal database structures that might be employed within corporate efforts such as Google. Because of its widespread crowdsourced nature, OSM tends to be more detailed than other maps for less well-served communities such as people cycling or walking, and its metadata is richer, as they are created by people who are intimately familiar with the areas that they are mapping. A map by users for users. 

Next up was Dr Rachel Sippy, a Research Associate with the Department of Genetics who presented how geospatial data factored into epidemiological research. In her work, the questions of ‘who’, ‘when’, and ‘where’ a disease outbreak occurred are important, at it is the where that gives her research a geographical focus. Maps, however, are often not detailed enough to provide information about an outbreak of disease among a population or community as maps can only mark out the incident site, the place, whereas the spatial context of that place, which she denotes as space, is equally as important in understanding disease outbreaks.  

Of ‘Space’ and ‘Place’: Rachel at the November 2023 Data Champion forum

It can be difficult, however, to understand what a researcher is measuring and what types of data can be used to measure space and/or place. Spatial data, as Rachel pointed out, can be difficult to work with and the researcher has to decide if spatial data is a burden or fundamental to the understanding of a disease outbreak in a particular setting. Rachel discussed several aspects of spatial data which she has considered in her research such as visualisation techniques, data sources and methods of analysis. They all come with their own sets of challenges and researchers have to navigate them to decide how best to tell the fundamental story that answers the research question. This essentially comes down to an act of curation of spatial data, as Rachel pointed out, quoting Mark Monmoneir, that “not only is it easy to lie with maps, it’s essential”. In doing so, researchers working with spatial data would have to navigate the political and cultural hierarchies that are explicitly and implicitly inherent to places, and any ethical considerations relating to both the human and non-human (animal) inhabitants of those geographical locations. Ultimately, how data owners choose to model the spatial data will affect the analysis of the research, and with it, its utility for public health. 

After lunch, both Martin and Rachel sat together to hold a combined Q&A session and a discussion emerged around the topic of subjectivity. A question was raised to Rachel regarding mapping and subjectivity, as it was noticed that how she described place, which included socio-cultural meanings and personal preferences of the inhabitants of the place, can be considered to be subjective in manner. Rachel agreed and alluded back to her presentation, where she mentioned that these aspects of mapping can get fuzzy as researchers would have to deal with matters relating to identity, political affiliations and personal opinions, such as how safe an individual may feel in a particular place. Martin added that with the OSM project the data must be objective as possible, yet the maps themselves are subjective views of objective data.  

Rachel and Martin answering questions from the Data Champions at the November 2023 forum

Martin also brought to attention that maps are contested spaces because spaces can be political in nature. Rachel added that sometimes, maps do not appropriately represent the contested nature of her field sites, which she only learned through time on the field. In this way, context is very important for “real mapping”. As an example, Martin discussed his “UK collision data” map, created outside the University, which states where collisions have happened, giving the example of one of central Cambridge’s busiest streets, Mill Road: without contextual information such as what time these collisions occurred, what vehicles were involved, and the environmental conditions at the time of the accident, a collision map may not be that valuable. To this end, it was asked whether ethnographic research could provide useful data in the act of mapping and the speakers agreed. 

The Data Picture

I was recently named one of “the next generation of [library] leaders” as part of the CILIP 125, having been recognised as an individual who contributes energy and knowledge to improving and impacting their organisation. My area of expertise, and thus recognition, lies with the use of data within libraries. As a data analyst for the Office of Scholarly Communications at Cambridge University Library, my role focuses on empowering decisions with data driven understanding – such as supporting the Springer Nature negotiations. To develop my understanding of data, and its role within a wider organisation, further, I engage with data beyond the library – such as the Big Data London conference and the Carruthers and Jackson Data Leaders’ Summer School. Reflecting on the use of data in the wider world, what can be expected of the library and data?


The summer school provided practical advice, proven methodologies, and guidance that could apply across a variety of businesses. The course is designed to provide insight on the workflow of data officers, and their role within an organisation – no matter its stage of data maturity and literacy. Over the course of the ten weeks, leading experts discussed the role of a chief data officer (CDO), both as a business development opportunity, and as a career path for individuals. It explored the risk and governance of data within an organisation, and the final weeks focused strongly on the role of people and teams associated with data.

Peter Jackson and Caroline Carruthers addressed the differing types of CDO and described a pendulum between ‘risk aversion’ and ‘value added’. Understanding the balance between secure and proper data governance (GDPR for example) and providing value through data (such as setting up automation). The pendulum of risk to reward is relevant to many roles, including those within the library. Understanding the need to divide time and energy between creating policies and getting decision making results, is just as relevant to my role as a chief data officer. In my role I have supported decision making staff through data production, but equally, to instil a culture of data, time and energy must be dedicated to risk aversion, through tasks of researching data management, preparing training sessions for data storage, and supporting staff in data preparation.

Another important concept introduced was the DIKW pyramid – Data, Information, Knowledge, Wisdom – for understanding the value created from data. The base of the pyramid is (raw) Data, which can be processed into (useful) Information. This Information is data with meaning and a purpose and can be organised into (insightful) knowledge. Knowledge combines experiences, values, insights, and contextual information, which can then transcend to (integral) Wisdom. Wisdom is considered a deeper understanding with ethical implications and the ability to define ‘why’. The DIKW pyramid provided a frame of thought for presenting and approaching future data projects. Understanding the requirement to provide, data, information or knowledge, to better support a decision-making team.

To develop communication skills, expert Scott Taylor, known as The Data Whisperer, spoke about the three V’s for data storytelling: Vocabulary, Voice and Vision. Combining an accessible vocabulary, with a common voice will illuminate the business vision, and why that is important. This overarching concept for an organisations data approach can be scaled down to support individual data workers, to provide value – which should either grow, improve or protect the business case. Understanding how to communicate the data is a key skill as “Hardware comes and goes, software comes and goes, but data remains”. And that data that remains should be used to either grow, improve or protect the business, such that data gathered should be usable data!

At Big Data London, the organisation Women in Data hosted conversations about nurturing a culture of learning within data teams. Pulling from their experiences from minority backgrounds, the speakers highlighted the power in upskilling, sharing skills across teams and being an advocate on oneself and skills. As for what to upskill, data literacy was a hot topic across the conference. Data literacy, also called data fluency and data confidence, is the combination of ability, skills and confidence surround data and its uses. Data literacy enables more efficient work, and begs the question, what is the base level of data literacy / confidence across the library? Librarians use data daily; checking in/out material, answering students’ queries, or tracking the use of space, but are all librarians confident to use that data? This is an area I hope to explore further at the CUL, to ensure staff can use the data they have to support decisions.


Engaging with the world of data provides a big picture of the possibilities within the library. Conversations of AI (Artificial Intelligence), data policies and maturity, and shiny-new databases, software, and services, demonstrate the growing adoption of data, and therefore, libraries should follow suit. Actively taking snippets of larger conversations, developing ideas within the library space, and exploring the possibilities with data will help libraries thrive in this world of technological growth.


Should the UK make a deal with Springer Nature?

This is a guest post by Prof. Stephen J. Eglen on the concurrent negotiations between the UK academic sector and the publisher Springer Nature. Prof. Eglen is a Fellow of Magdalene College and Professor of Computational Neuroscience in the Department of Applied Mathematics and Theoretical Physics at the University of Cambridge. This post does not necessarily reflect the view of Cambridge University Libraries.

The UK academic sector is currently in discussion with Springer Nature around a renewed ‘read and publish’ deal for journal content. I understand that most institutions are likely to reject the current deal, but wish to continue negotiations. My position is that further discussions with Springer Nature are futile; we should stop accepting ‘transformative deals’. The likely effect of this deal would be that more of Springer Nature’s content may be openly available to read, but with the ‘paywall’ shifted to the publish side. Here I list my key objections:

  1. There is still no justification for the high APCs (9500 EUR + taxes) for Nature tier journals. Accepting a deal, regardless of the level of discounts that could be achieved, is implicitly accepting their business model. Springer Nature declined to engage with the Journal Comparison Service run by cOAlition S that aims to help understand how costs are determined.
  2. Springer Nature’s view is that ‘gold OA’ is the only viable way to open access. Other models for open access are available, and show promise, including diamond OA journals and Subscribe to Open. However, Springer Nature assert that “they haven’t found a way of making them financially sustainable”.  If we accept a gold-only view of open access,  how can we objectively assess the sustainability of alternative models?
  3. A move to a ‘gold only’ OA world would shift the barrier from reading to publishing content. Springer Nature recently announced a waiver policy for researchers from about 70 lower income countries. This still excludes many researchers worldwide e.g. from Brazil and South Africa, perpetuating neo-colonial attitudes towards the creation of scholarly content and reinforcing existing institutional inequalities within countries. Any waiver programme for APCs should be “no-questions-asked” regardless of where researchers are based. This would need to be properly costed and part of the justification of the APC (point 1).
  4. As of January 2023, several UK institutions have rights retention policies in place, with more expected to follow in the coming months. Individual researchers can also use rights retention strategy by themselves. Rights retention statements allow researchers to meet UK funder’s requirement by depositing their author-accepted manuscript without embargo. I believe Springer Nature should publicly state that they will allow any author worldwide to maintain their rights on their own author-accepted manuscripts.
  5. Over half of Springer Nature’s hybrid journals failed to meet their 2021 targets for open access articles within hybrid journals.  Those hybrid journals that fail again this year to meet their targets will be removed from cOAlition S’s transformative journal program.  Having some journals ineligible for cOAlition S funding but part of a UK read-and-publish deal would further complicate an already confusing system.  It would also question Springer Nature’s commitment to open access.

A detailed public critique of the deal is not possible because of the confidential nature of the negotiations.  Finances aside, I feel there was one element that was simply unworkable and unethical due to it requiring scholars to keep one aspect confidential if the deal were accepted.

The UK is one of only a few countries with a  heavy reliance on transformative agreements.  Sweden has already decided that transformative agreements are not sustainable and the transition period should finish at the end of 2024. Coalition S has also confirmed it will end its support of hybrid journals by the end of 2024. I would like to see the UK move away from transformative agreements. We could instead work internationally to promote more ethical and sustainable alternatives that put scholars at the heart of scholarly communication. In particular, the APC model has been tried, and introduces as many headaches as it has tried to solve. 

It is time instead to try new approaches.  There are several interesting models being developed by forward-looking organizations that the UK could endorse.  For example, MIT press recently launched shift+OPEN as a way to flip subscription based journals to diamond open access model.  Another interesting approach is Subscribe to Open where journals drop their paywall if a threshold amount of subscriptions are received.  Money saved on dealing with legacy publishers like Springer Nature is better spent investing in our own infrastructure and new approaches.