Showing posts with label digital preservation. Show all posts
Showing posts with label digital preservation. Show all posts

Thursday, 29 November 2018

World Digital Preservation Day 2018


Today is World Digital Preservation Day.  It is the day when the digital preservation community around the world come together to celebrate the collections that have been preserved, the access has been maintained, and the work that is being done to preserve our digital legacy.

Organised by the Digital Preservation Coalition and supported by digital preservation networks all over the world, World Digital Preservation Day raises awareness of the strategic, cultural and technological issues which make up the digital preservation challenge. Since the first public website was launched nearly 30 years ago, there has been an explosion of digital content worldwide. This ‘born digital’ content is tomorrow’s cultural heritage, and it’s our job to ensure that we collect and preserve this digital history for future generations.

When I first heard about this day I was excited.  Ever since I attended the iPRES (International Conference on Digital Preservation) conference in 2014 when it was held in Melbourne, I have been fascinated, passionate, intrigued with digital preservation.  Unfortunately it isn't something that is shared with my institution.  Digital preservation costs money, sometimes a lot of money.  And that is something that my institution is very loathe to part with.  (I should know, I've been campaigning for a new institutional repository system since 2010 with no luck, and that is peanuts compared to a preservation system).

But, in my own way, I have been steadily pushing the preservation agenda, one step at a time.

So, many people don't really understand what digital preservation actually is.  Essentially it is the coordinated and ongoing set of processes and activities that ensure long-term, error-free storage of digital information with a means for retrieval and interpretation for the entire time span the information is required.  Preservation isn't just digitisation, however digitisationis part of preservation.

So with that in mind, this is what we have been doing in this space for the last few years.

Institutional repository (IR)
Our IR actually does preservation reasonably well (which is surprising as not much else works with it).  It produces a PREMIS datastream, which is the international standard for metadata to support the preservation of digital objects and ensure their long-term usability.  If only we understood how the IR platform actually uses it.....  Our IR also does versioning very well.  While not strictly 'digital preservation', it is a form of preservation as each version is maintained in the system (metadata and documents). However this is where our investment into digital preservation ends.  Our documents are at best stored as standard PDFs, at worst Word documents or other proprietary file formats.  And we do not actively maintain the files in the system, but rather just 'believe' that they will work when we next try to download one.

Digital collections
Our institution has a number of digital research collections and the degree of "preservation-ness" for each one varies.  One has PDF/A as a standard for any scanned documents, another has straight PDFs, and the third (a collection of images) has RAW, TIFF and JPG files.  However, like the IR, the files are not actively maintained for fixity and access.  This is a new focus area for us (and for the system that we use) and I'm hoping that it will further develop in this space in the future.

Research data
Unfortunately, beyond backup, no digital preservation activities are performed on our research data.  Like musch of our research infrastructure we are working with sub-standard ad hoc systems.  When we can better manage it from an administrative aspect, then hopefully we will be able to better manage the digital preservation of it.


Digital preservation is something that I get very excited about. However I'm not naive enough to realise that it can and probably is a very dry subject to many people. So, just to make it a bit more interesting, here are some websites that I think are pretty cool.

The Museum of Obsolete Media has over 500 current and obsolete physical media formats covering audio, video, film and data storage.

And the ‘Bit List’ of Digitally Endangered Species, a crowd-sourcing list of which digital materials the community things are most at risk.

But perhaps the coolest thing about World Digital Preservation Day is that it gives us a chance to have cake!

Happy World Digital Preservation Day everyone!

Thursday, 1 November 2018

CAUL Research Repository Days 2019

The 2018 CAUL Research Repository Days were held in Melbourne over 29-30 October. Although there was much discussion over many different topics, the program was very much focused on interoperability between systems which is a trend that I have observed the IR community heading towards. With a well running platform, repository work is less about the 'publication' and more about how systems interact with each other.

The below is a summary of some of the themes that were of particular interest to my institution and myself. 
 

CAUL Projects

Review of Australian Repository Infrastructure Project

Much of Day 1 was in discussion of FAIR. Drafted in 2014 but published in 2016, the FAIR principles (of Findable, Accessible, Interoperable and Reusable) are a set of 14 metrics designed to determine the level of FAIRness of an output or system. In response to this, CAUL proposed a project in 2017 to determine how improvements to repository infrastructure can be made across the sector to increase the FAIRness of Australian-funded research outputs. The final report has just been released.

The project followed seven project working groups designed to examine the current repository infrastructure, international repository infrastructure developments, repository user stories, ideal state for Australian repository infrastructure, next generation repository tools, and make recommendations for the possible "Research Australia" collection of research outputs. The first six group findings are included in the report, while the seventh, the "Research Australia" recommendations, is due at the end of 2018.

Each working group provided a report of their findings. Most were not surprising and were generally what we have known to be the case for some time. In summary (and in no particular order), they include:
  • Although nine institutions had new generation repository software, many of the others had ageing infrastructure that perhaps had not been able to be funded since the ASHER funding in 2007, with VITAL particularly mentioned for dropping in number 
  • Ageing software was identified as a weakness of repository infrastructure, as was the lack of automation and identifiers 
  • Research outputs were the most common output in IRs, followed by theses and research data. Other output types included archival collections, journal, images and course materials 
  • Institutions numbers were almost equal in terms of having an OA policy, a statement or partial policy, or no OA policy at all 
  • Only 5 institutions supported research activity identifiers (although they didn’t specify RAiDs in particular) 
  • 13 institutions had a digital preservation strategy for the IR content, with a further 3 developing a strategy 
  • Most successful initiatives have stable secure funding. 
  • Recommendation that CAUL seek consortia membership of COAR 
  • List of general repository requirements. 

The seventh group is looking at the feasibility of a "Research Australia" portal as a single-entry point to a collection of all Australian Research outputs. This is similar to the RUN proposal some years ago. Views were extremely mixed regarding this. Responses included that it is duplicating what we already have with Google Scholar and TROVE, whether it would be OA or metadata only, the quality of metadata, and questions over unique institutional requirements. Three possibilities have been proposed - upgrade TROVE to provide all necessary reporting needs, develop a new portal harvesting repositories (similar to the OpenAIRE model), or developed a shared infrastructure. 
 

Collecting and Reporting of Article Processing Charges (APCs)

Another CAUL project currently underway is the APC project determining the cost of article processing charges for institutions. Several options are proposed. Less preferred include creating a fund code in the finance system of the institution or querying the finance system using a selection of keywords. Another less preferred option is obtaining reports from publishers or making them provide this information as part of the subscription agreement. What is likely to be proposed is a very manual method of extracting a dataset from Web of Science, Scopus and Dimensions, either by institution or nationally, run it against the unPaywall API to find which are OA publications, deduping on DOI then, using the publisher list price for APCs, determining the cost of the APC payment based on the corresponding author institution. A couple of institutions have done this calculation internally with varying results. My own use of the unPaywall API has shown it to be unreliable in terms of finding OA outputs as false positives can be returned, however it seems to be the most promising tool to date in this respect. 
 

Retaining Rights to Research Publications

A survey of Australian university IP policies has been undertaken to identify potential barriers to the implementation of a national licence in Australia, similar to the UK-SCL licence. The key element of the UK-SCL licence is to retain the right to make the accepted manuscript of scholarly articles available publicly for non-commercial use (CC BY NC 4.0) from the moment of first publication. An embargo can be requested (by either the author or the publisher) for up to 12 months. However only 13 Australian universities have an IP policy that would be supportive of this licence. Recommended as the next step by CAUL is to approach Universities Australia for consideration and the development of guidelines for alignment of IP policies. 
 

Statement on Open Scholarship

A final CAUL project is the Statement on Open Scholarship which is a call to action around advocacy, training, publishing, infrastructure, content acquisition and education resources. The review period ends at the end of October. 
 

FAIR Data


Natasha Simons, ARDC, reported on an American Geophysical Union project designed to enable FAIR data. The project objectives were to look at FAIR-aligned repositories and FAIR-aligned publishers. There is a push for repositories to be the home for data rather than the supplementary section of journals. A commitment statement has been produced with a set of criteria that repositories must meet in order to enable FAIR data. (USC can meet about half of the requirements with the current infrastructure and policies).

In terms of Australian repositories, the AGU project may influence subsequent projects in other research disciplines. As publishers are moving away from data in supplementary sections of journals to data in (largely domain) repositories, trusted repositories (the Core Trust Seal) are becoming increasingly important.

Ginny Barbour, AOASG, proposed a new acronym, "PID+L" (pronounced, piddle) as the essential minimum of metadata required for research outputs to be FAIR:
  • PID 
    • ORCID 
    • DOI for all outputs 
    • PURL for grants 
  • Licence (machine readable) 
Note that we are unable do this with our current infrastructure. 

ORCID


Simon Huggard, Chair of the ORCID Advisory Group, provided a snapshot of the ORCID Consortium in Australia. There are 41 organisations that are part of the consortium with 32 integrations completed (by 29 consortium members). The most popular system used for integration are custom integrations, followed by Symplectic, Pure, Converis, IRMA, Scholar One and ViVo. Seven institutions have done full ORCID authentication integration so that researchers can sign into ORCID using their institutional credentials. Currently there are 90K Australian researchers registered with an ORCID, up from 30K at the beginning of 2016.

The ORCID Consortium has developed a Vision 2020 which aims to have all active researchers in Australia with an ORCID, and all using their ORCID throughout the research lifecycle. The ARC and NHMRC will integrate ORCID into their grant management systems (which they have done, and which will be live in the next couple of weeks), and where possible, government agencies to draw upon ORCID data for research performance reporting and assessment.

There are challenges in integrating ORCID institutionally, most common being private profiles (early profiles were set to private by default) and synchronisation issues, particularly duplicates where metadata may be slightly different in varying source data. Another challenge is getting ORCID to be displayed in IRs. When asked about this, the ARC replied that although this is a requirement of their OA mandate, at present it is not a problem although it will be in the future.

Digital preservation


Jaye Weatherburn, University of Melbourne, gave a keynote presentation on digital preservation and the role that libraries, in particular IRs, need to play in this. Digital preservation is a series of managed activities necessary to ensure continued access to digital materials for as long as necessary. There are several reasons for looking at digital preservation - decay of storage media, rapidly advancing technology leading to obsolescence, fragility of digital materials, and protection against corruption and accidental deletion. A digital preservation strategy can be used to monitor these risks. Long term preservation however is not a 'set and forget'. It is an iterative process to ensure the life of a document is maintained. Without digital preservation there is no access to materials in the long term.

It should be noted that our IR doesn’t ‘do’ digital preservation beyond saving PDF files of outputs where available, along with metadata. The FIA collection does digital preservation slightly better, in that the PDF/A standard is used for master representations. While the Herbarium perhaps does it the best, with RAW, TIFF and JPG files being saved for each image. However, without a digital preservation system such as Rosetta, we are not so much preserving our digital data but rather just backing it up to protect against deletion.

Closely aligned with this theme of preservation is that of trustworthiness of a repository (which also includes the organisation). There are two frameworks that are commonly used for examining the trustworthiness of repositories - the Core Trust Seal, and the Audit and Certification of Trustworthy Digital Repositories based on ISO16363. Both can be self-assessed and provide a good means of documenting gaps, although the Core Trust Seal is less intensive on resourcing and time. This is something that I have been keen to do for USC since I first heard about it at the iPRES conference in 2014 and is something I will complete once a decision is made regarding a future system.

Below is a word-cloud of what attendees thought digital preservation meant to them:
 

Other interesting things: 

Idea of incentivising scholarly communication via cryptocurrency.

Chris Berg, RMIT, opened with a keynote on blockchains as a tool to govern the creation of knowledge. Blockchains are economic infrastructure on which new forms of social organisation can be built. Chris states that academic publishing is a subset of a general problem that has afflicted publishing and the knowledge economy since the invention of the internet. The RMIT Blockchain Innovation Hub project had the idea of incentivising scholarly communication via cryptocurrency - a token to pay and reward for peer review, sharing citations, reading, etc. In terms of economic modelling, journal publishing can be viewed as a 'club'. The aim of the project was to bring transparency to the peer review process, provide digital copyright authentication and verification, and to provide incentives and rewards for the different aspects of the scholarly communication lifecycle. Enter 'JournalCoin'… Subscriptions, article processing fees and peer reviewers could be paid by JournalCoin, and rewards for such things as fast peer reviews, formatting, royalties, rankings and citations paid via JournalCoin. The journal is then the platform upon which the incentives are paid. 

IRUS-UK pilot in Australia

CAVAL is currently running a project on implementing IRUS-UK in Australia. IRUS (Institutional Repository Usage Statistics) started in the UK in 2012 and sought to provide a standards-based service with auditable usage data. The aim was to reduce duplication of effort by IR managers and present a uniform set of usage data regardless of the IR platform. IRUS data is COUNTER-compliant. IRUS-UK now does this for about 140 IRs in the UK. A pilot has been underway in Australia involving University of Melbourne, Victoria University, University of Queensland, University of Sydney and Monash University to evaluate the usefulness of IRUS in Australia. Several of these institutions reported on their experience, which was largely positive. One advantage of the IRUS statistics is that they exclude 'false positive' metrics, resulting in slightly lower statistics than the native IR ones. CAVAL reported that if usage of IRUS goes ahead, maximum benefit will be realised if the majority of Australian universities participate and individual universities will be able to benchmark against each other. 

Social Media campaigns

Susan Boulton, GU, provided an outline on a pilot the Library ran to promote their IR through social media. By using national/international events (such as World Malaria Day, Sustainability Week, and Dementia Month), blog posts and social media mentions were written showcasing the research that was in their IR. To prepare time was spent planning, sourcing open access content, identifying champion event owners, and preparing the social media material. These small social media events provided a significant jump in IR traffic and downloads. Another benefit was the improved relationship between researchers and the Library, as researchers can see another value-added service.

Tuesday, 21 October 2014

Digital Preservation


Digital Preservation



I had the opportunity to attend the 11th iPRES conference held in Melbourne - the first time the conference had been held in the Southern Hemisphere! The digital preservation community is relatively small so the conference, with 177 delegates, was well attended and included staff from university libraries, state and national libraries, archives, museums, commercial vendors and technology developers. 46% of delegates were international, giving a wide range of expertise and experiences. As a novice to the digital preservation space there was much to take in, however a number of ‘themes’ were apparent as was indicated by the various initiatives in the sector.


Why care about research data management?


Although slightly outside the 'scope' of digital preservation, research data management is still an important part of the process. Without properly managed data, there is nothing for us to preserve. Ross Wilkinson, Director of the Australian National Data Service (ANDS), gave a keynote presentation on the reasons why researchers, and institutions, need to embrace research data management. These include compliance with the Australian Code of Responsible Conduct of Research and funding bodies, as well as the long term management of researcher data. Additionally, sharing data can lead to data citation and increased collaboration. Citation rates, particularly if the collaboration is international, can increase by up to three times.

Institutions also need to embrace research data management, and need to start viewing research data as a research output rather than just a research by-product. Sharing data and making data available also supports the research ambitions of the institution, which in turn supports the research strategy of the institution. The Vice Chancellor of the University of Tasmania, Professor Peter Rathjen, has said that as reputation is very important to research institutions, and as libraries make substantial contributions to that reputation, libraries (being the experts on digital collections) should be supported in creating world class data collections which in turn can help an institutions reputation.


Changes to the scholarly communication model


Andrew Treloar, Director of Technology for ANDS, gave a presentation on the changes to the scholarly communication process. The existing/previous system of registration (journal submission), certification (peer review), awareness (discovery services) and archiving (libraries, publishers, archiving services such as Portico) is changing. Much more of the scholarly process is now on the web and is wholly digital, and includes not just publications but also datasets, slides, wikis, processes, workflows, and logs, all packaged up and surrounded by metadata. This change of process has meant that the existing systems of scholarly communication have also shifted. Registration systems now include things such as Protein Banks and Wiki Pathways, certification systems include open peer review, awareness systems include open wikis and e-lab books, and archiving systems include institutional and data repositories (which although technically are not "preservation" systems, ANDS recognised that they are as good as we have at the moment).

However this new scholarly communication system does pose a problem for citing sources. In the past, the majority of publications remained ‘findable’ but now cited webpages can change and cited datasets may disappear. Common web platforms are increasingly used for scholarship, such as wikis, Github, Twitter and Wordpress. Many of these have desirable characteristics such as versioning, time stamping and social embedding, but they record rather than archive. This is a problem as they capture critical elements of the scholarly record which will be lost over time.

There is a difference between the scholarly process (which is short term, write many/read many, no guarantees provided) to the scholarly record (longer term, write once/read many, attempt to provide guarantees). We need to start thinking about moving from recording the scholarly process to archiving the scholarly record.

Another presentation by Herbert Van de Sompel, Los Alamos National Laboratory, talked about problems with referencing web content. At the same time that we are adding things to the web, we are losing as well. A report into social media documentation following the Egyptian Revolution in 2011 found that 10% of references had disappeared off the web a year after the event.

There are two problems with referencing web content - link rot, where links stop working, and content drift, where linked content changes over time. A study shortly to be published by PLoS found that 15% of links in articles submitted in 2012 were already dead, with 35% being dead after 5 years.

There are ways the preserve this content. An experimental solution for Zotero will automatically send a website to the internet archives when it is bookmarked in Zotero. The user then gets a link to the website as well as a link to the web archive version along with the date in Zotero. Other solutions are identifier systems such as DOIs, however these are only as good as the agency or organisation that is managing them.


Preservation processes



Preservation Policies

"Without a policy framework a digital library is little more than a container for content". Preservation policies have important features, one of which is to inform various stakeholders of the digital archives and provide transparency about the approaches to preservation. It also enhances the "trustworthiness" of the archive. However most institutions either do not have a preservation policy or it is so out of date that it is effectively useless.

Barbara Sierman, National Library Netherlands, reported on the European SCAPE project which looked at, amongst other things, what is required for a preservation policy. Using the few resources available, they developed a catalogue of policy elements as a guideline to help improve preservation policies. A maturity level model was also introduced to score how mature the policy framework for an institution is.


Collection Profiling

Maureen Pennock, Head of Digital Preservation at the British Library, outlined their process for collection profiling, which is a precursor to collection preservation. It is important as it defines the type of content being held and what needs to be preserved. It is also important for looking at what shouldn't be preserved – disposal is just as important as preservation.

When collection profiling, a number of factors are examined, including a summary of the content, acquisition methods and formats, preservation intent, issues with preservation, and any sub-collections and different representations of collections. An example was given using the British Library web archives. Once collection profiling is completed, file format assessment and preservation profiling can begin.


File Format Assessment

The concept of file format endangerment and obsolescence is important when considering digital preservation activities. According to Heather Ryan, Assistant Professor at University of Denver, file format endangerment describes the possibility that information stored on a particular file format will not be interpretable or renderable using standard methods within a certain timeframe, whereas file obsolescence occurs when information stored in a particular file format is no longer accessible using current technologies.

File format assessments analyse the risk of the use of a file format. The British Library has analysed a number of file formats, looking at such characteristics as development status, adoption, software support, complexity, external dependencies, technical protection mechanisms and legal issues. These are used in conjunction with other preservation activities to determine the endangerment level of particular file formats. Examples were given using TIFF, JP2 and PDF. Researchers at the University of Denver have also identified three key factors when considering file endangerment: availability of rendering software, specifications and community/third party support.


Technical Registries

Technical registries are used in digital preservation to enable organisations to maintain definitions of the formats, format properties, software, migration pathways, etc., needed to preserve content over the long term. There are numerous technical registries around at the moment, with more currently being developed. The most commonly used registries are PRONOM, Freebase and DROID.

One of the problems with most registries is that they are fixed information models, and are difficult to evolve. All registries also do not describe the same thing, with different use cases and technical requirements for each registry. Each institution needs to evaluate their own requirements and risk analysis of digital preservation practices.

Two new registries have been created by Preservica and the National and State Libraries of Australasia (NSLA). Preservica has released their Linked File Format Registry along with their new version of software. Through peer-to-peer collaboration, the registry can be added and edited, with each installation choosing which changes the wish to incorporate into their own instance of the registry – a first of its kind ‘linked’ data registry. The NSLA is also funding a project to create a Digital Preservation Technical Registry to collate the information for the various registries into one place.

There is interest in automating as much of the digital preservation workflow as possible. The Data Archiving and Networked Services in the Netherlands has developed a system called Epimenides that can check whether a newly ingested file is in an acceptable or preferred format, check whether it is migratable to another format, and do any migrations necessary automatically depending on rules set up in the system.


Preservation and digital repositories


Tomasz Miksa, from SBA Research Austria, presented on risk assessments for digital repositories. Digital preservation of a repository system should not only concern the content, but also the workflows, software, and metadata. If the system cannot visualise the data or see the data, then the repository is useless. Technical aspects as well as organisational needs should be considered during risk assessments. Repository systems may be required to undergo several digital preservation actions in order to preserve both the system and workflows. Dependence on external services and insufficient documentation are dependencies for digital preservation actions.

The Danish State and University Library reported on a self-assessment of their digital repository that they undertook in order to expose the drivers for digital preservation, improve staff and management understanding of the digital preservation challenges, and to enable benchmarking with other digital preservation organisations.


Conclusions


Digital preservation is still largely ad hoc in many institutions, with confusion regarding exactly just what it is. Collaboration amongst the digital preservation community was a key message throughout the conference, although as one delegate mentioned, there are currently multiple file format registries being worked on, so while everyone seemingly agrees that collaboration is the key, it remains to be seen to be put into practice.

And for those institutions that are currently facing an uphill battle with management to try to show the value of digital preservation, the suggestion was made to try masking the records or files in an organisation that are older than 5 years old as unreadable, and see what sort of outcry there is when people try to access them.

Image credit: http://lts2.evault.com/homepage/digital-preservation/