Showing posts with label ORCID. Show all posts
Showing posts with label ORCID. Show all posts

Thursday, 1 November 2018

CAUL Research Repository Days 2019

The 2018 CAUL Research Repository Days were held in Melbourne over 29-30 October. Although there was much discussion over many different topics, the program was very much focused on interoperability between systems which is a trend that I have observed the IR community heading towards. With a well running platform, repository work is less about the 'publication' and more about how systems interact with each other.

The below is a summary of some of the themes that were of particular interest to my institution and myself. 
 

CAUL Projects

Review of Australian Repository Infrastructure Project

Much of Day 1 was in discussion of FAIR. Drafted in 2014 but published in 2016, the FAIR principles (of Findable, Accessible, Interoperable and Reusable) are a set of 14 metrics designed to determine the level of FAIRness of an output or system. In response to this, CAUL proposed a project in 2017 to determine how improvements to repository infrastructure can be made across the sector to increase the FAIRness of Australian-funded research outputs. The final report has just been released.

The project followed seven project working groups designed to examine the current repository infrastructure, international repository infrastructure developments, repository user stories, ideal state for Australian repository infrastructure, next generation repository tools, and make recommendations for the possible "Research Australia" collection of research outputs. The first six group findings are included in the report, while the seventh, the "Research Australia" recommendations, is due at the end of 2018.

Each working group provided a report of their findings. Most were not surprising and were generally what we have known to be the case for some time. In summary (and in no particular order), they include:
  • Although nine institutions had new generation repository software, many of the others had ageing infrastructure that perhaps had not been able to be funded since the ASHER funding in 2007, with VITAL particularly mentioned for dropping in number 
  • Ageing software was identified as a weakness of repository infrastructure, as was the lack of automation and identifiers 
  • Research outputs were the most common output in IRs, followed by theses and research data. Other output types included archival collections, journal, images and course materials 
  • Institutions numbers were almost equal in terms of having an OA policy, a statement or partial policy, or no OA policy at all 
  • Only 5 institutions supported research activity identifiers (although they didn’t specify RAiDs in particular) 
  • 13 institutions had a digital preservation strategy for the IR content, with a further 3 developing a strategy 
  • Most successful initiatives have stable secure funding. 
  • Recommendation that CAUL seek consortia membership of COAR 
  • List of general repository requirements. 

The seventh group is looking at the feasibility of a "Research Australia" portal as a single-entry point to a collection of all Australian Research outputs. This is similar to the RUN proposal some years ago. Views were extremely mixed regarding this. Responses included that it is duplicating what we already have with Google Scholar and TROVE, whether it would be OA or metadata only, the quality of metadata, and questions over unique institutional requirements. Three possibilities have been proposed - upgrade TROVE to provide all necessary reporting needs, develop a new portal harvesting repositories (similar to the OpenAIRE model), or developed a shared infrastructure. 
 

Collecting and Reporting of Article Processing Charges (APCs)

Another CAUL project currently underway is the APC project determining the cost of article processing charges for institutions. Several options are proposed. Less preferred include creating a fund code in the finance system of the institution or querying the finance system using a selection of keywords. Another less preferred option is obtaining reports from publishers or making them provide this information as part of the subscription agreement. What is likely to be proposed is a very manual method of extracting a dataset from Web of Science, Scopus and Dimensions, either by institution or nationally, run it against the unPaywall API to find which are OA publications, deduping on DOI then, using the publisher list price for APCs, determining the cost of the APC payment based on the corresponding author institution. A couple of institutions have done this calculation internally with varying results. My own use of the unPaywall API has shown it to be unreliable in terms of finding OA outputs as false positives can be returned, however it seems to be the most promising tool to date in this respect. 
 

Retaining Rights to Research Publications

A survey of Australian university IP policies has been undertaken to identify potential barriers to the implementation of a national licence in Australia, similar to the UK-SCL licence. The key element of the UK-SCL licence is to retain the right to make the accepted manuscript of scholarly articles available publicly for non-commercial use (CC BY NC 4.0) from the moment of first publication. An embargo can be requested (by either the author or the publisher) for up to 12 months. However only 13 Australian universities have an IP policy that would be supportive of this licence. Recommended as the next step by CAUL is to approach Universities Australia for consideration and the development of guidelines for alignment of IP policies. 
 

Statement on Open Scholarship

A final CAUL project is the Statement on Open Scholarship which is a call to action around advocacy, training, publishing, infrastructure, content acquisition and education resources. The review period ends at the end of October. 
 

FAIR Data


Natasha Simons, ARDC, reported on an American Geophysical Union project designed to enable FAIR data. The project objectives were to look at FAIR-aligned repositories and FAIR-aligned publishers. There is a push for repositories to be the home for data rather than the supplementary section of journals. A commitment statement has been produced with a set of criteria that repositories must meet in order to enable FAIR data. (USC can meet about half of the requirements with the current infrastructure and policies).

In terms of Australian repositories, the AGU project may influence subsequent projects in other research disciplines. As publishers are moving away from data in supplementary sections of journals to data in (largely domain) repositories, trusted repositories (the Core Trust Seal) are becoming increasingly important.

Ginny Barbour, AOASG, proposed a new acronym, "PID+L" (pronounced, piddle) as the essential minimum of metadata required for research outputs to be FAIR:
  • PID 
    • ORCID 
    • DOI for all outputs 
    • PURL for grants 
  • Licence (machine readable) 
Note that we are unable do this with our current infrastructure. 

ORCID


Simon Huggard, Chair of the ORCID Advisory Group, provided a snapshot of the ORCID Consortium in Australia. There are 41 organisations that are part of the consortium with 32 integrations completed (by 29 consortium members). The most popular system used for integration are custom integrations, followed by Symplectic, Pure, Converis, IRMA, Scholar One and ViVo. Seven institutions have done full ORCID authentication integration so that researchers can sign into ORCID using their institutional credentials. Currently there are 90K Australian researchers registered with an ORCID, up from 30K at the beginning of 2016.

The ORCID Consortium has developed a Vision 2020 which aims to have all active researchers in Australia with an ORCID, and all using their ORCID throughout the research lifecycle. The ARC and NHMRC will integrate ORCID into their grant management systems (which they have done, and which will be live in the next couple of weeks), and where possible, government agencies to draw upon ORCID data for research performance reporting and assessment.

There are challenges in integrating ORCID institutionally, most common being private profiles (early profiles were set to private by default) and synchronisation issues, particularly duplicates where metadata may be slightly different in varying source data. Another challenge is getting ORCID to be displayed in IRs. When asked about this, the ARC replied that although this is a requirement of their OA mandate, at present it is not a problem although it will be in the future.

Digital preservation


Jaye Weatherburn, University of Melbourne, gave a keynote presentation on digital preservation and the role that libraries, in particular IRs, need to play in this. Digital preservation is a series of managed activities necessary to ensure continued access to digital materials for as long as necessary. There are several reasons for looking at digital preservation - decay of storage media, rapidly advancing technology leading to obsolescence, fragility of digital materials, and protection against corruption and accidental deletion. A digital preservation strategy can be used to monitor these risks. Long term preservation however is not a 'set and forget'. It is an iterative process to ensure the life of a document is maintained. Without digital preservation there is no access to materials in the long term.

It should be noted that our IR doesn’t ‘do’ digital preservation beyond saving PDF files of outputs where available, along with metadata. The FIA collection does digital preservation slightly better, in that the PDF/A standard is used for master representations. While the Herbarium perhaps does it the best, with RAW, TIFF and JPG files being saved for each image. However, without a digital preservation system such as Rosetta, we are not so much preserving our digital data but rather just backing it up to protect against deletion.

Closely aligned with this theme of preservation is that of trustworthiness of a repository (which also includes the organisation). There are two frameworks that are commonly used for examining the trustworthiness of repositories - the Core Trust Seal, and the Audit and Certification of Trustworthy Digital Repositories based on ISO16363. Both can be self-assessed and provide a good means of documenting gaps, although the Core Trust Seal is less intensive on resourcing and time. This is something that I have been keen to do for USC since I first heard about it at the iPRES conference in 2014 and is something I will complete once a decision is made regarding a future system.

Below is a word-cloud of what attendees thought digital preservation meant to them:
 

Other interesting things: 

Idea of incentivising scholarly communication via cryptocurrency.

Chris Berg, RMIT, opened with a keynote on blockchains as a tool to govern the creation of knowledge. Blockchains are economic infrastructure on which new forms of social organisation can be built. Chris states that academic publishing is a subset of a general problem that has afflicted publishing and the knowledge economy since the invention of the internet. The RMIT Blockchain Innovation Hub project had the idea of incentivising scholarly communication via cryptocurrency - a token to pay and reward for peer review, sharing citations, reading, etc. In terms of economic modelling, journal publishing can be viewed as a 'club'. The aim of the project was to bring transparency to the peer review process, provide digital copyright authentication and verification, and to provide incentives and rewards for the different aspects of the scholarly communication lifecycle. Enter 'JournalCoin'… Subscriptions, article processing fees and peer reviewers could be paid by JournalCoin, and rewards for such things as fast peer reviews, formatting, royalties, rankings and citations paid via JournalCoin. The journal is then the platform upon which the incentives are paid. 

IRUS-UK pilot in Australia

CAVAL is currently running a project on implementing IRUS-UK in Australia. IRUS (Institutional Repository Usage Statistics) started in the UK in 2012 and sought to provide a standards-based service with auditable usage data. The aim was to reduce duplication of effort by IR managers and present a uniform set of usage data regardless of the IR platform. IRUS data is COUNTER-compliant. IRUS-UK now does this for about 140 IRs in the UK. A pilot has been underway in Australia involving University of Melbourne, Victoria University, University of Queensland, University of Sydney and Monash University to evaluate the usefulness of IRUS in Australia. Several of these institutions reported on their experience, which was largely positive. One advantage of the IRUS statistics is that they exclude 'false positive' metrics, resulting in slightly lower statistics than the native IR ones. CAVAL reported that if usage of IRUS goes ahead, maximum benefit will be realised if the majority of Australian universities participate and individual universities will be able to benchmark against each other. 

Social Media campaigns

Susan Boulton, GU, provided an outline on a pilot the Library ran to promote their IR through social media. By using national/international events (such as World Malaria Day, Sustainability Week, and Dementia Month), blog posts and social media mentions were written showcasing the research that was in their IR. To prepare time was spent planning, sourcing open access content, identifying champion event owners, and preparing the social media material. These small social media events provided a significant jump in IR traffic and downloads. Another benefit was the improved relationship between researchers and the Library, as researchers can see another value-added service.

Saturday, 22 July 2017

Your online scholarly identity...why is it important?


There are so many tools out there now to help researchers build their online identity – ORCID, Google Scholar, ResearchGate, Academia.edu, Twitter, Facebook, LinkedIn, and the list goes on. But what do we mean by ‘online identity’ and why is it so important to researchers? (A side question could also be why we, as librarians, care about this when obviously, many researchers don’t? But that is another story altogether).

So, what do we mean by online identity. Everyone has an online identity if you have had anything to do with the internet. Type your name into Google and most people will get at least one hit on their name with some people getting many hits. But why is this important? Your online identity is what people who don’t know you (but know about you) search for – prospective employers, colleagues, rivals, people you have just met at a conference. From a professional viewpoint, it is important that these people find your information, and most importantly, find accurate information easily.

Maintaining and curating your professional identity is the key to achieving this. But in a minefield of options which ones do you choose?

You have probably been using products that you have either been using in a previous life, or that your colleagues are already using. However, no matter which product (or products) you use, it is critical that these be kept up to date. There is nothing worse that searching and finding someone’s profile only to find that the most recent publication listed is from four years ago, or where they say they work is clearly not the case.

Where I work, we have three that we recommend that researchers keep up to date – ORCID, Google Scholar profile and of course, the institutional repository.

[Note: for the purpose of this I am discounting social media such as Twitter, Facebook and LinkedIn – all of which are important in disseminating information about yourself and your works.]



ORCID has quickly risen to be a universal identifier for scholarly outputs, and, why wouldn’t it? It is free to use, independent, and has a fantastic developing team behind it. Being a truly independent identifier also means that it easily integrates with external systems (unless you work where I do, where nothing really seems to integrate with anything else – but that is a whole other issue). With ORCID you can upload your outputs, or connect to another data source (such as Scopus or Research Data Australia) and any of your works listed in these data sources will be imported into your profile. You can add grant information, education and work details, as well as provide links to other online profiles that you maintain. By sharing your ORCID iD URL with colleagues you can provide a quick, one-stop-shop to all professional information about yourself.
The second online profile we recommend curating is a Google Scholar My Citations Profile. Google Scholar is big….not as big as Google, but in academic circles it is still pretty awesome. Having a Google Scholar profile set up is one key way for other people to find your research. Of course, it relies on your publications being indexed and available in Google, although there is the option of manually adding selected metadata about those that aren’t in Google. A word of warning – it is very easy to accidently add publications that aren’t yours to your profile, or for someone else to add your publications to their profile. Careful checking and periodic searches of your works may be advisable.


Which brings us to the old favourite, the institutional repository. Love it or hate it, it is here to stay and likely tied into your institutions reporting requirements. So, love it or hate it, you may be required to use it. I love our institutional repository, even the funny quirks that make it frustrating to work with, but it serves a purpose and a function in our academic community. And it is easy for our researchers. All they need to do is send in the metadata of their recent publications and the Research Collections team does the rest! It is a way for researchers to archive their outputs, preserve them, make them open access, and provides a nice link that they can send out to colleagues. It is also indexed by Google, so readily discoverable for anyone searching for your publications. I can’t speak highly enough for institutional repositories around the world – they are by far the best way to promote your research.
This then leaves us with the ‘badies’ of the online scholarly identity world – ResearchGate and Adademia.edu. I must admit that each of these have their place in the scholarly identity world, and all three of them are wildly popular in various disciplines. My reasons for disliking them are purely selfish – they are a rival to our institutional repository. And in this day and age where everyone is time poor, why invest your energies in keeping these up to date as well as the other critical profiles. ResearchGate and Academica.edu are both proprietary products that could turn off at any time. They have no preservation strategy and no commitment to keeping your work safe. If you wish to use them to disseminate your works, then please, please, please do so as an additional method to those listed above.



So now that you have all methods for people to find information about you and your works, what do you do? The key is to spend some time each week, month, couple of months (depending on your frequency of publication) updating them. As stated above, there is nothing worse than colleagues or other interested people finding an out of date profile. By keeping these updated, you are presenting your best self to those that want to find you.





Wednesday, 28 June 2017

Notes from CAUL Repository Community Day 2017

The annual CAUL Repository Community Day was held last Monday as a satellite event alongside the Open Repositories 2017 conference in Brisbane.  It was an excellent exchange of ideas and knowledge (as always) with lots of institutions doing many wonderful things in their institutional repository space.

Program: here 

Kathleen Shearer (COAR) started the day with a wonderful talk about COAR (Confederation of Open Access Repositories) and the problem with the current scholarly publishing system.  As library journal subscriptions continue to rise, Kathleen posed the question that if we are collecting published content and putting it into our repositories, are we then just perpetuating a flawed system?  In addition, the need researchers feel to be published in "luxury" journals with high impact factors is forcing some researchers to work in 'trendy' areas, rather than doing the more important less glamorous work that needs to be done.  Kathleen used the example of the recent Zika virus outbreak.  Prior to the outbreak those researching in this area had problems getting their articles published in top journals.  However as soon as the virus made it to the US, it became a 'trendy' topic and suddenly researchers could publish anywhere.  This model of selective, elitist journal publishing is something that we need to move away from, although in today's academic climate of performance measurement (both internally in our institutions, and externally by the government) being based on citation rates and visibility in a particular subscription database, it will be a long time before we change.

COAR is also investigating the so-called 'next generation' repositories, and what such a repository may look like. As new technologies and services are developed, we need to strive to continue to make our repositories relevant.  This is something that I feel really strongly about.  The answer to this, according to Kathleen, is to strengthen and add value to our repository networks.  Repositories are critical to our future visions of libraries and serve two roles: to showcase and provide access to the scholarly record of our own institutions; and as nodes in a global knowledge commons.  To support this global nature of repositories, COAR launched the Aligning Repository Networks International Accord on 8th May 2017.  A shared vision of this strategic coordination will facilitate data exchange such as cross regional harvesting between networks and repositories.  The problem is that Australia doesn't have a formal 'repository network', something that the Australasian Repository Working Group (ARWG) is examining. Australia also lacks a national aggregator to exchange information with international aggregators such as OpenAIRE.

Another shared vision of the Accord is interoperability - common vocabularies and metadata guidelines.  This is also a focus of the ARWG, and is seen as a core feature of repositories.  There are however many challenges in improving interoperability with so many disparate systems and 'business-purposes' for our repositories.  The ARWP sees interoperability requiring collaboration and a common understanding between repositories in Australia, as well as globally.  Part of this common understanding is providing a consistent approach to metadata standards and vocabularies   The NISO 'free_to_read' and 'licence_ref' tags are a beginning as this will help to identify open access content across systems.  University of New England and Deakin University are the pioneers in this area, having implemented these tags in their institutional repositories already.  

In so much as we have a national aggregator, TROVE is especially important to Australian repositories in that it harvests content into a single database.  Due to the number of disparate systems in Australia, there is much variation of the quality of the data going into TROVE, with a large variety of metadata schemas and formats.  Julia Hickie from TROVE spoke about the sort of data that is going into TROVE from our repositories.  Identifiers, in particular, have proliferated in the last five years.  In spite of this, ORCIDs are nearly invisible in the TROVE data, accounting for only about 1% of harvested records.  This indicates that very few repositories are actually recording the ORCID identifier in their metadata records.  Having said this, there are now over 17,000 ORCIDs in TROVE, up from 2000 two years ago.  In order to improve the quality of the data in TROVE, Julia advised that it is a important to do a 'health check' on your repository data that is being harvested by TROVE every now and then, and particularly if you change something or move to a new system.  Things to look at are:
  • make sure the repository URL is in a dc:identifier
  • author ORCIDS are URLs in dc:relation
  • ARC/NHMRC grant identifiers as a URL in dc:relation, and in the format http://dx.doi.org/[doi]
  • DOIs are URLS in dc:relation or dc:identifier
  • open access indicator uses free_to_read
  • Creative Commons licence information in full URL form in dc:rights or ali:licence_ref
  • rights statements are in full URL form in dc:rights or ali:licence_ref.
Check the OAI-PMH feed yourself to ensure that it is working correctly. 

There are a number of institutions doing fantastic things with their repositories.  Some of these presented on the day include:
  • Robin Burgess from the University of Sydney spoke about collecting non-traditional research outputs (NTROs) and bridging the gap between these outputs and traditional outputs, which had previously been kept in separate repositories.  Consultation with academics showed that a repository that would suit these "defiant objects", or the "rule breakers", had to have a visually rich interface, with many academics already having these outputs showcased on personal websites.  The system had to be able to display these outputs along with the ephemeral information that accompanied them.
  • Janice Chan from Curtin University reported on their recent move to Dspace using an external vendor, Atmire.  They took the opportunity during the migration to change some work practices, including reducing the organisational structure in the system from many faculties to just two groups (research papers and theses, although they plan to add a third for grey literature) and to improve their metadata standards with the addition of funder and rights fields.
  • Kate Croker from University of Western Australia talked about their Repository Project, whereby they enriched their Pure repository with grant data, researcher profiles, publication collections, and the like.   The project provided scope for extensive collaboration between stakeholders and external departments, strengthening the relationship in the process.  Through this collaboration, a shared vision for the repository was produced which has helped to shape the direction, as well as build support, for the repository.  [Kate also has presented this at ALIA Online - paper and presentation available here]
  • Bernadette Houghton from Deakin University spoke on using Omeka software for digital collections.  She especially mentioned some recommendations surrounding the use of third party plugins versus those available on Omeka.org.  

Bernedette Houghton also spoke about self assessment for repositories against ISO 16363 which Deakin University completed in 2013.  This assesses features such as governance, technical infrastructure, security and preservation.  In completing the project Deakin University looked at the existing literature, performed the self assessment, and addressed the areas of improvement.  Bernedette provided a list of recommendations for anyone wanting to do a repository self assessment, including:
  • choose a tool (it doesn't have to be ISO 13636)
  • review the criteria from the start
  • understand ISO 13636's conceptual nature (based on OAIS)
  • preference local knowledge over ISO suggested documentation
  • allocate resources to address areas of improvement.
[Note: I first heard about this at iPRES 2014, and have been wanting to do this for our repository ever since, however I know how dismally ours would fail so haven't had the nerve until we move to a new system.  More from Bernedette can be found here  http://dx.doi.org/10.1045/march2015-houghton]

The day finished up with a general discussion on how institutions are dealing/coping with ERA and the absence of publications for HERDC (and if anything had taken it's place).  Most (all?) institutions are still collecting publications on an annual basis, whether as ERA prepation, internal reporting, KPI's, or for "just-in-case" the government changes it's mind about HERDC.  There have also been many institutions where the responsibility for publication collection has shifted from the research office to the library.  

There was some discussion of the Engagment and Impact Assessment and how institutions went about completing this.  Mary-Anne Marrington from University of Queensland reported on their method.  
[Note: Our Office of Research decided not to participate in the pilot (held in 2017), so I can't comment on our process.]

The definition of "open access" for ERA (and more generally), in particular the difference between open access and free access, and which was correct for ERA, produced some lively debate.  Different institutions have different definitions, and would require some clear guidelines.  The ARC Open Access Policy is due out soon (the draft has been out for some months), so hopefully, at least for ERA purposes, this will be much clearer.


There were others that presented, and their presentations were fantastic.  It is always good (although a bit deperssing) to hear what others are doing in this space.  I always leave with many good ideas that I would love to implement back at home, but as always, resourcing (staffing and money) are always a problem.  So, many of these ideas remain in limbo.   I think it would be good for the powers that be to take note of Kathleen Shearer's comment, "repositories are a technology and technologies change".  We need to continue to strive to make our repositories relevant in this changing landscape, and continue to value add services in either the repository layer or in the network layer above.

Friday, 19 February 2016

Easy as 1, 2, 3...

I know what I am interested in, and I know what I am passionate about, but seldom do these coincide with things that others are interested in and passionate about.  Except ORCID.

ORCID (Open Researcher and Contributor Identifier) is something that seems to resonate with a whole bunch of people, from hard core researchers to administrators and librarians.  For those that are saying "but what do flowers have to do with researchers?", ORCID is a way to disambiguate researchers, especially those with similar sounding names.  By giving each researcher a unique number, they can then go and 'tag' their research publications, data, grants, and many other research outputs as theirs.  It is like the grand-daddy of researcher identifiers.  And it has landed in a big way.

However, I get the feeling that ORCID is more popular with research administrators than with the researchers themselves.  There are a multitude of reasons why research institutions can benefit from ORCID (streamlining processes, reporting on research undertaken,identifying research resulting from grants awarded to staff, etc), but the benefits to researchers are not as obvious.  Sure, being about to differentiate between the various "Tim Smiths" that work at the institution would be nice, but what else?  What is there that drives the researcher to maintain their ORCID profile?

I recently read an article by The Research Whisperer that sums this up nicely, and I created a sketch note about it.  And I must say, speaking as someone that rarely gets a like or retweet on Twitter, this has gone galactic!  It has definitely hit a chord with many people who work either in research or around research.  It is a credit to the author of the article (Jonathan O'Donnell) for writing such a wonderful piece.

Some of the comments I have received via Twitter include:
@BecOwen74, Just made my day. Thanks! (from @jod999)
 I love this summary of ORCID from the Research Whisperer. This image makes its uses very straighforward (from @Ashley_UQL)
Get your research & profile out there - for all academics, postdocs, PhD students.  Love the graphic! (from @LareenNewman) 
And to top it off, the author even included it in his blog post on The Research Whisperer!

How chuffed am I!!

So, without further ado, here is my sketch note.  I hope you enjoy.