Featured
Linked Data, DOI, RDF, and Dublin Core
On a quest to reach the holy grail of the Semantic Web.
What started as a straightforward and elegant model by Eric Miller as the Resource Description Framework (RDF) has become incredibly complicated. When I first came upon this model, in the early days of the Dublin Core standard discussions, it was a framework. That is, it was a self describing way to transmit an XML file. The RDF formed a wrapper around the XML by including the XML Schema (DTD) so that whoever received the file would know what the elements, attributes, allowed ranges, etc. were, and be able to use the data included without further hunting for file descriptions, translations of the fields (elements), how they related, etc. RDF has grown up and become embroiled in discussions of triples, Subject-Object-Predicate discussions, and its use as the basis for linked data and even the basis for the final real implementation of the Semantic Web.
Read MoreUsing a ‘Collabulary’ to Create a Taxonomy
Libraries and librarians have been the gatekeepers to knowledge stores for more than 200 years. As their collections grew, they invented ways to easily find the information and knowledge they stored by creating classification systems and then subject headings to identify the concepts or topics represented in the items being stored. Every major language now has at least one classification system, and most countries have created and adopted classification and subject access systems, such as the Universal Decimal System (UDC) or Lenin’s outline of knowledge for Russia. In the United States, the use of the Dewey Decimal Classification system, Sears Subject Headings, and the Library of Congress Classification system is widespread.
Read MoreBreaking Down Automatic Metadata Generation/Extraction
There are two approaches to automatic metadata generation/extraction and within those, many variations. The first is statistical. Generally speaking, the types are Bayesian, vector, neural nets, automatic clustering, etc. These methods work off the principle that if in a big set of data two words occur together frequently then those words are related conceptually.
Read MoreTaxonomies — Here, There, Everywhere
I went to one of my favorite bookstores the other day. I have been awaiting the release of E. E. Knight’s Winter Duty. Looking at the overhead signs for sections, I realized that this is the beginning of a taxonomy….My realization that the bookstore is using a taxonomy made me realize that taxonomies are merely a way of grouping information together in a way that is relevant to the user.
Read MoreCalculating the ROI of Semantic Enrichment
We are often asked for help in calculating the potential return on investment (ROI) of investing in a taxonomy. How this calculation is done depends on the type of organization and who will be the users of the taxonomy. For any content-intensive organization, using a taxonomy increases findability, which in turn leads to greater utility and value for the organization’s content assets. Within the enterprise, this translates into hours and dollars saved by reducing the time required to find things, as well as the time and expense of redoing research or other work that can’t be found.
Read MoreThe Race for Relevancy: Information Foraging
The key word in Web 2.0 is relevancy. The surge in social networking sites and mobile internet means we are always connected and always feeding information into the greater whole. Easy access to relevant information is emerging as a major focus as the internet continues to be flooded with exponentially more information.
Read MoreStandards: Straightjacket or Straight Edge?
Standards were not popular with would-be railroad tycoons in the 1880s. A standard set of railroad tracks eliminated some profitable loading and unloading work. Incompatible tracks kept the other guys’ rolling stock on another guy’s tracks. Eventually standards emerged because a couple of tycoons figured out how to make even more money by standardizing. Many industries go through this wide open approach only to find that over time certain methods or designs become part of the furniture of living.
Read MoreLost In Translation
Different cognitive paths are driven by the language we speak and our culture drives the way we organize and categorize information. These differences will lead us to very different discoveries and conclusions.
Read MoreField by Any Other Name
July 26, 2010 — In one of the chatty podcasts about software engineering, one of the commentators pointed out that metadata was one of the most critical aspects of information governance. We agree. The chatter continued and another podcast voice mentioned that there was a relationship between database field names and metadata. I realized that…
Read MoreMicrosoft and Its Taxonomy of Intents
July 19, 2010 — Xiaosin Yin and Sarthak Shah wrote a paper published in April 2010, “Building Taxonomy of Web Search Intents for Name Entity Queries.” We revisited this paper because Yahoo announced that it was considering a service that used popular queries for a new Yahoo news service. You can read a summary of…
Read More