Access Innovations, Inc. Announces Release of the Semantic Fingerprinting Web Service Extension for Data Harmony Version 3.9

Access Innovations, Inc. announces the Semantic Fingerprinting Web service extension as part of their Data Harmony Version 3.9 release. Semantic Fingerprinting is a managed Web service offered to scholarly publishers to disambiguate author names and affiliations by leveraging semantic metadata within an existing publishing pipeline.

Read More

Putting Human Intelligence To Work To Enhance the Value of Information Assets

Semantic enhancement extends beyond journal article indexing, though the ability of users to easily find all the relevant articles (your assets) when searching still remains the central purpose. Now, in addition to articles, semantic “fingerprinting” is used for identifying and clustering ancillary published resources, media, events, authors, members or subscribers, and industry experts. The system…

Read More

The Size of Your Thesaurus

During the initial stages of discussing a new taxonomy project, I am frequently asked questions like: How granular does my taxonomy need to be? How many levels deep should the vocabulary go? And especially: How many terms should my thesaurus have? The answer is—of course—it depends.

Read More

Access Innovations, Inc. Announces Release of the Smart Submit Extension Module to Data Harmony Version 3.9

Access Innovations, Inc. announces the Author Submit extension module as part of their Data Harmony Version 3.9 release. Author Submit is a Data Harmony application for integration of author-selected subject metadata information into a publishing workflow during the paper submissions or upload process. Author Submit facilitates the addition of taxonomy terms by the author. With Author Submit, each author provides subject metadata from the publisher taxonomy to accompany the item they are submitting. During the submission process, Data Harmony’s M.A.I. core application suggests subject terms based on a controlled vocabulary, and the author chooses appropriate terms to describe the content of their document, thus enabling early categorization and selection of peer reviewers and support for trend analysis.

Read More

Identifying the Contributors

Collaboration is a key component of research. Original research papers with a single author are — particularly in the life sciences — a vanishing breed. This makes it difficult to identify author contributions and acknowledgements, as well as to mine any data from the unstructured information.

Read More

Hold the Mayo! A study in ambiguity

When we (at least those of us in Greater Mexico) hear of or read about Cinco de Mayo there is no question in our minds that “Mayo” refers to the month of May. The preceding “Cinco de” (Spanish for “Fifth of”) pretty much clinches it. Of course, if the overall content is in Spanish, there might still might be some ambiguity about whether it is the holiday that is being referred to, or simply a date that happens to be the one after the fourth of May. (As in “Hey, what day do we get off work?” “The fourth of July, I think.”)

Read More

The Supposed Advantages of Statistical Indexing

I would like to make some observations about statistics-based categorization and search, and about the advantages that their proponents claim.

First of all, statistics-based co-occurrence approaches do have their place. For wide-ranging bodies of text such as email archives and social media exchanges, and for assessing the nature of an unknown collection of documents, a well-defined collection of concepts covering a pre-determined area of study and practice is not possible. For lack of this foundation, and for lack of other practical approaches, attempts at analysis fall back on less-than -ideal mathematical approaches.

Read More