Friday, 19 February 2016

Easy as 1, 2, 3...

I know what I am interested in, and I know what I am passionate about, but seldom do these coincide with things that others are interested in and passionate about.  Except ORCID.

ORCID (Open Researcher and Contributor Identifier) is something that seems to resonate with a whole bunch of people, from hard core researchers to administrators and librarians.  For those that are saying "but what do flowers have to do with researchers?", ORCID is a way to disambiguate researchers, especially those with similar sounding names.  By giving each researcher a unique number, they can then go and 'tag' their research publications, data, grants, and many other research outputs as theirs.  It is like the grand-daddy of researcher identifiers.  And it has landed in a big way.

However, I get the feeling that ORCID is more popular with research administrators than with the researchers themselves.  There are a multitude of reasons why research institutions can benefit from ORCID (streamlining processes, reporting on research undertaken,identifying research resulting from grants awarded to staff, etc), but the benefits to researchers are not as obvious.  Sure, being about to differentiate between the various "Tim Smiths" that work at the institution would be nice, but what else?  What is there that drives the researcher to maintain their ORCID profile?

I recently read an article by The Research Whisperer that sums this up nicely, and I created a sketch note about it.  And I must say, speaking as someone that rarely gets a like or retweet on Twitter, this has gone galactic!  It has definitely hit a chord with many people who work either in research or around research.  It is a credit to the author of the article (Jonathan O'Donnell) for writing such a wonderful piece.

Some of the comments I have received via Twitter include:
@BecOwen74, Just made my day. Thanks! (from @jod999)
 I love this summary of ORCID from the Research Whisperer. This image makes its uses very straighforward (from @Ashley_UQL)
Get your research & profile out there - for all academics, postdocs, PhD students.  Love the graphic! (from @LareenNewman) 
And to top it off, the author even included it in his blog post on The Research Whisperer!

How chuffed am I!!

So, without further ado, here is my sketch note.  I hope you enjoy.


Tuesday, 12 January 2016

Open Access in scholarly communication

Open Access (OA) has come along way since the idea was formalised in 2002 as the Budapest Open Access Initiative.  It has since become the catch call for scholarly communication, providing an ideal tool in the dissemination of research results and publications far and wide.

Since it's inception, there are many models of implementing OA - gold, hybrid, delayed and green.  Green OA is the preferred, at least from this Repository Manager's point of view, however I fully appreciate that some researchers do not want a less than perfect version of their work out on display to the world.  This is where the Gold/Hybrid/Delayed route comes into play - still all very legitimate OA options.

No matter which OA method is chosen, the most important thing is that research results are made available to anyone that is interested, regardless of access to subscription library databases, and that universities and other research institutions recognise the importance in providing the infrastructure and the means to make this research accessible.


(Based on the article by Mohammad Reza Ghane (2014) Open Access Policy. International Journal of Information Science and Management)

Thursday, 7 January 2016

Impact of research on society

As I experiment in the sketch noting world, I use as my test topic an excellent article from academics at Charles Stuart University.

Societal impact has come under intense discussion lately in Australia as the Government prepares to trial an Assessment and Impact Framework from 2018, with a pilot to be run in 2017.  This will be the first time that institutions nationwide will be involved in an assessment of this nature, which is touted to be along the lines of the UK Research Excellence Framework (REF).

But what is societal impact?  According to the Australian Research Council (ARC), impact is
"the demonstrable contribution that research makes to the economy, society, culture, national security, public policy or service, health, the environment, or quality of life, beyond contributions to academia" (ARC Research Impact Principles and Framework).
This has started many conversations by worried university administrators as to how such impact can be measured.

This is where this article, and my naive efforts at sketch noting, helps us to understand.  Bracing for Impact: The role of information science in supporting societal impact, by Lisa Given, Wade Kelly and Rebekah Willson, was presented at the ASIST 2015 Annual Meeting held in the United States.


Tuesday, 27 October 2015

Researcher identifiers

Researcher identifiers.....these unique sets of characters that identify a particular researcher as themselves, removing any ambiguity with other researchers of a similar name, are so important to the career of a researcher.  Not only the researcher though, but also the institution that needs to report on many different metrics relating to research output.

For those that do not know what a researcher identifier is or what the advantages are, here is a quick rundown.

However I am constantly amazed at the lack of care factor that researchers show towards researcher identifiers.  Don't get me wrong, some researchers "get it" and appreciate the importance of these characters.  These are the ones that actively maintain their publications and add any missing ones.  I love these researchers.  But the other 90%....I just don't get it!  I don't understand why they do not invest in the process.

The problem with Scopus:
Scopus is slightly different from the rest in that it is a system generated number assigned by Elsevier.  When a publisher sends metadata to Scopus, a fancy algorithm tries to match the author based on name spelling, format and affiliation.  If there is insufficient evidence to match with an existing author in the system, Scopus will automatically create a new one.  You can see how it is very easy for authors to end up with multiple identifiers in Scopus.  We recently did a "data cleansing" exercise in Scopus to try to identify and de-dupe multiple Scopus identifiers.  I contacted each author to explain the situation and provided step by step instructions on what to do (I thought it was good that the author engages with the process so that they may then learn and keep on top of it themselves in the future).  The most I found was one author that had 13 different identifies!! When you think about how much their research impact metrics were diluted out by having publications spread over 13 different identifiers, the mind boggles.

So why do researchers just not care?  Is it that they are too busy, the process to complicated or do they just not understand the importance?  Or maybe they just don't even know about researcher identifiers - no one has told them?

At USC we are trying to address the lack of care-factor with regards to researcher identifiers.  A comprehensive online guide has been produced and is regularly updated (however the limitations of our website mean that it is not very discoverable).  The librarians, when talking to researchers about outputs and metrics mention it.  And we have even run a competition during our recent USC Research Week conference for a chance to win a $100 voucher for every researcher identifier reported to the Library.  Emails to new staff ask if they have any researcher identifiers (from which we can obtain publications metadata for entry into our institutional repository, the USC Research Bank).

During the recent USC Research Week conference, the Library had a display encouraging researchers to think about their online research profile and what they can do to improve it.  One of the most contentious posters was a "Top 10 @ USC" which listed the top 10 authors with, amongst other things, citations in Scopus, Web of Science and Google Scholar.  The bottom line is that if a researcher doesn't have a researcher identifier or has a poorly managed researcher identifier then their publications will not be able to be measured using conventional recognised metrics.

Perhaps an addendum to the Top 10 @ USC poster is to put "We really struggled to find you because you didn't have a researcher identifier".



eResearch Australasia 2015

Last week I attended what is one of the best conferences of the year - eResearch Australasia, held in Brisbane.  It is always a very inspiring event and I always come back filled with ideas to put into practice at work.

This year was a bit different to previous years in that it had a more library/human capital focus.  Previous years have been heavy on the technology side which, while interesting, was often slightly over my head.  This year is different.

The main themes prevalent throughout the conference were:

  • The importance of libraries and librarians for open data
  • Linked open data, not just shared data
  • Data as an institutional asset
  • The connected researcher.

Next year is being held in Melbourne during October - only a year to implement all my ideas before the next round.

Friday, 26 June 2015

Research data sharing


Sharing research data is increasingly becoming more popular, and while not synonymous with traditional scholarly publishing yet, it is nevertheless moving in that direction. We, as an aspiring research institution, need to start thinking about depositing and sharing “publications and data”, rather than treating research data as a special entity, if in fact we treat it as anything at all.

There are a number of benefits to the institution and the researcher for sharing data. Demonstrating good practice and research integrity raises the profile of the university and individual researcher. Sharing data makes it citable, which in turn can lead to increased citation metrics for both the publications associated with the data and the data itself.  This is a good thing.  Increased exposure from the data records can help foster new collaborations in research areas not previously thought of. And funding opportunities may improve due to a healthier research ecosystem and greater integration between systems and researcher profiles.

When reading about the positives for data sharing it is hard to understand why there is such a resistance to sharing within academic circles. Do researchers fear they will not be recognised or credited for their data? If the data has a good framework around it making it easy to obtain, understand and cite, then this risk should be reduced. Or do they fear “getting scooped”? Embargoing the data may be the solution to this.

Often institutions and policy makers have a perception that it is the “big” data that needs the most help when it comes to managing and sharing. This is usually not the case. Big data often has a more robust framework surrounding the collection and management of it – often due to requirements of funding organisations. The problem is with small data – the multitude of small spreadsheets that researchers maintain, often without adequate management, code keys, storage, backup… If data is managed correctly during the collection and analysis stage, it makes it all the easier for sharing once work has been completed. Data that is managed correctly – i.e. has a good framework around it – is more likely to be used and therefore cited. Unfortunately, citations are the name of the game in order to stay current in research.

For every risk or concern that researchers or institutions can throw up for sharing, there will always be a solution. Data should be shareable. Publically funded data should definitely always be shareable. The risk to institutions for not sharing data – non-compliance with policy and funding agreements, reputational damage, poor practice, low awareness –means that institutions should lead by example and facilitate the sharing infrastructure.

Sometimes there are legitimate concerns about sharing data – it is identifiable, confidential, private? What is the best way to manage this sort of data? Is it shareable? In these cases, the metadata can be available with mediated access to the data. When it comes to data of a sensitive nature, there will always need to be someone that can respond to requests.

Research data is an institutional asset, and as such should be treated as such. Unrecognised effort is a prime precursor to disengagement from researchers, staff and the community. And as an asset, you (whether the researcher, lab technician, administrator, executive, institution) should be treating research data with the respect it deserves.

“Products of research are not just publications” – NSF senior policy specialist Beth Strausser.

Graphic: http://d7.library.gatech.edu/research-data/home

Tuesday, 21 October 2014

Digital Preservation


Digital Preservation



I had the opportunity to attend the 11th iPRES conference held in Melbourne - the first time the conference had been held in the Southern Hemisphere! The digital preservation community is relatively small so the conference, with 177 delegates, was well attended and included staff from university libraries, state and national libraries, archives, museums, commercial vendors and technology developers. 46% of delegates were international, giving a wide range of expertise and experiences. As a novice to the digital preservation space there was much to take in, however a number of ‘themes’ were apparent as was indicated by the various initiatives in the sector.


Why care about research data management?


Although slightly outside the 'scope' of digital preservation, research data management is still an important part of the process. Without properly managed data, there is nothing for us to preserve. Ross Wilkinson, Director of the Australian National Data Service (ANDS), gave a keynote presentation on the reasons why researchers, and institutions, need to embrace research data management. These include compliance with the Australian Code of Responsible Conduct of Research and funding bodies, as well as the long term management of researcher data. Additionally, sharing data can lead to data citation and increased collaboration. Citation rates, particularly if the collaboration is international, can increase by up to three times.

Institutions also need to embrace research data management, and need to start viewing research data as a research output rather than just a research by-product. Sharing data and making data available also supports the research ambitions of the institution, which in turn supports the research strategy of the institution. The Vice Chancellor of the University of Tasmania, Professor Peter Rathjen, has said that as reputation is very important to research institutions, and as libraries make substantial contributions to that reputation, libraries (being the experts on digital collections) should be supported in creating world class data collections which in turn can help an institutions reputation.


Changes to the scholarly communication model


Andrew Treloar, Director of Technology for ANDS, gave a presentation on the changes to the scholarly communication process. The existing/previous system of registration (journal submission), certification (peer review), awareness (discovery services) and archiving (libraries, publishers, archiving services such as Portico) is changing. Much more of the scholarly process is now on the web and is wholly digital, and includes not just publications but also datasets, slides, wikis, processes, workflows, and logs, all packaged up and surrounded by metadata. This change of process has meant that the existing systems of scholarly communication have also shifted. Registration systems now include things such as Protein Banks and Wiki Pathways, certification systems include open peer review, awareness systems include open wikis and e-lab books, and archiving systems include institutional and data repositories (which although technically are not "preservation" systems, ANDS recognised that they are as good as we have at the moment).

However this new scholarly communication system does pose a problem for citing sources. In the past, the majority of publications remained ‘findable’ but now cited webpages can change and cited datasets may disappear. Common web platforms are increasingly used for scholarship, such as wikis, Github, Twitter and Wordpress. Many of these have desirable characteristics such as versioning, time stamping and social embedding, but they record rather than archive. This is a problem as they capture critical elements of the scholarly record which will be lost over time.

There is a difference between the scholarly process (which is short term, write many/read many, no guarantees provided) to the scholarly record (longer term, write once/read many, attempt to provide guarantees). We need to start thinking about moving from recording the scholarly process to archiving the scholarly record.

Another presentation by Herbert Van de Sompel, Los Alamos National Laboratory, talked about problems with referencing web content. At the same time that we are adding things to the web, we are losing as well. A report into social media documentation following the Egyptian Revolution in 2011 found that 10% of references had disappeared off the web a year after the event.

There are two problems with referencing web content - link rot, where links stop working, and content drift, where linked content changes over time. A study shortly to be published by PLoS found that 15% of links in articles submitted in 2012 were already dead, with 35% being dead after 5 years.

There are ways the preserve this content. An experimental solution for Zotero will automatically send a website to the internet archives when it is bookmarked in Zotero. The user then gets a link to the website as well as a link to the web archive version along with the date in Zotero. Other solutions are identifier systems such as DOIs, however these are only as good as the agency or organisation that is managing them.


Preservation processes



Preservation Policies

"Without a policy framework a digital library is little more than a container for content". Preservation policies have important features, one of which is to inform various stakeholders of the digital archives and provide transparency about the approaches to preservation. It also enhances the "trustworthiness" of the archive. However most institutions either do not have a preservation policy or it is so out of date that it is effectively useless.

Barbara Sierman, National Library Netherlands, reported on the European SCAPE project which looked at, amongst other things, what is required for a preservation policy. Using the few resources available, they developed a catalogue of policy elements as a guideline to help improve preservation policies. A maturity level model was also introduced to score how mature the policy framework for an institution is.


Collection Profiling

Maureen Pennock, Head of Digital Preservation at the British Library, outlined their process for collection profiling, which is a precursor to collection preservation. It is important as it defines the type of content being held and what needs to be preserved. It is also important for looking at what shouldn't be preserved – disposal is just as important as preservation.

When collection profiling, a number of factors are examined, including a summary of the content, acquisition methods and formats, preservation intent, issues with preservation, and any sub-collections and different representations of collections. An example was given using the British Library web archives. Once collection profiling is completed, file format assessment and preservation profiling can begin.


File Format Assessment

The concept of file format endangerment and obsolescence is important when considering digital preservation activities. According to Heather Ryan, Assistant Professor at University of Denver, file format endangerment describes the possibility that information stored on a particular file format will not be interpretable or renderable using standard methods within a certain timeframe, whereas file obsolescence occurs when information stored in a particular file format is no longer accessible using current technologies.

File format assessments analyse the risk of the use of a file format. The British Library has analysed a number of file formats, looking at such characteristics as development status, adoption, software support, complexity, external dependencies, technical protection mechanisms and legal issues. These are used in conjunction with other preservation activities to determine the endangerment level of particular file formats. Examples were given using TIFF, JP2 and PDF. Researchers at the University of Denver have also identified three key factors when considering file endangerment: availability of rendering software, specifications and community/third party support.


Technical Registries

Technical registries are used in digital preservation to enable organisations to maintain definitions of the formats, format properties, software, migration pathways, etc., needed to preserve content over the long term. There are numerous technical registries around at the moment, with more currently being developed. The most commonly used registries are PRONOM, Freebase and DROID.

One of the problems with most registries is that they are fixed information models, and are difficult to evolve. All registries also do not describe the same thing, with different use cases and technical requirements for each registry. Each institution needs to evaluate their own requirements and risk analysis of digital preservation practices.

Two new registries have been created by Preservica and the National and State Libraries of Australasia (NSLA). Preservica has released their Linked File Format Registry along with their new version of software. Through peer-to-peer collaboration, the registry can be added and edited, with each installation choosing which changes the wish to incorporate into their own instance of the registry – a first of its kind ‘linked’ data registry. The NSLA is also funding a project to create a Digital Preservation Technical Registry to collate the information for the various registries into one place.

There is interest in automating as much of the digital preservation workflow as possible. The Data Archiving and Networked Services in the Netherlands has developed a system called Epimenides that can check whether a newly ingested file is in an acceptable or preferred format, check whether it is migratable to another format, and do any migrations necessary automatically depending on rules set up in the system.


Preservation and digital repositories


Tomasz Miksa, from SBA Research Austria, presented on risk assessments for digital repositories. Digital preservation of a repository system should not only concern the content, but also the workflows, software, and metadata. If the system cannot visualise the data or see the data, then the repository is useless. Technical aspects as well as organisational needs should be considered during risk assessments. Repository systems may be required to undergo several digital preservation actions in order to preserve both the system and workflows. Dependence on external services and insufficient documentation are dependencies for digital preservation actions.

The Danish State and University Library reported on a self-assessment of their digital repository that they undertook in order to expose the drivers for digital preservation, improve staff and management understanding of the digital preservation challenges, and to enable benchmarking with other digital preservation organisations.


Conclusions


Digital preservation is still largely ad hoc in many institutions, with confusion regarding exactly just what it is. Collaboration amongst the digital preservation community was a key message throughout the conference, although as one delegate mentioned, there are currently multiple file format registries being worked on, so while everyone seemingly agrees that collaboration is the key, it remains to be seen to be put into practice.

And for those institutions that are currently facing an uphill battle with management to try to show the value of digital preservation, the suggestion was made to try masking the records or files in an organisation that are older than 5 years old as unreadable, and see what sort of outcry there is when people try to access them.

Image credit: http://lts2.evault.com/homepage/digital-preservation/