Access Insights
Pre or Post-Coordinate Indexing?
Most people think about what they want to search for and are willing to combine their concepts at the time of search. If you think of the way people search in Google, they put in the combination of terms they are thinking of; they are doing the coordination of terms. It is up to the search software to do the intersection of the terms for them and figure out the post-coordination.
Read MoreBreaking Down Automatic Metadata Generation/Extraction
There are two approaches to automatic metadata generation/extraction and within those, many variations. Generally speaking, the statistical approach types are Bayesian, vector, neural nets, automatic clustering, and the like. These methods work off the principle that if in a big set of data two words occur together frequently, then those words are related conceptually.atistical methods depend on the algorithms that the developer has come up with. Some vendors lock down the numerical recipes and provide few or no user controls. These systems usually require “training”. The way that they do this is to take a list of terms, often a thesaurus, taxonomy or authority file and a corpus of text. The system manager processes the inputs with spot checking and often manual intervention by a human subject matter expert. When the system has been trained, test queries are run to verify that the system is performing as desired on fresh content.
Read MoreSearch as Big Brother, Molding What You See and Think
A recent TED presentation is by Eli Pariser. He is the author of “The Filter Bubble: What the Internet Is Hiding From You.” A new and very interesting book. His talk is a synopsis of how the Google personalization algorithms effect search results. Google results are influenced by your own search history and other online activity. Any system such as Amazon, Yahoo, Bing ebay shopping systems depend heavily on personalization to serve you results. Traditional databases do not use profiles (yet) but they are often based on Verity, Vivisimo, Autonomy, Fast and other mathematically based search software so they could and they do serve up different results whenever the vectors are reset – that is every time additional data is added to the system with updates or metadata enrichment.
Read MoreHealth Care Logistics Chooses Metalogix for SharePoint Migration
Health Care Logistics has purchased Metalogix Migration Manager for SharePoint to migrate from SharePoint 2003 to SharePoint 2010. This opposed to using Microsoft tools, which would have taken much longer for the migration.
Read MoreAdventures of a TaxoTourist
The trip was awesome—a dream exotic vacation to Bali. It was not about eat, pray, love, but a rather unbalanced midpoint to meet my Oz-dwelling daughter. I enjoyed dashes of ecotourism and agritourism, but even in full vacation mode I couldn’t fully suppress my perspective as a taxonomist.
Read MoreThe Semantic Web Goes Mainstream
The recent MarkLogic User Conference was a watershed event for the publishers in attendance, many of whom are just beginning to strategize about the application of semantic technology to their content. After years of hearing “the Semantic Web is coming,” the message this time was that it’s no longer about “what” or “why,” but “how” publishers will leverage this technology. It has been 10 years since Tim Berners-Lee, Jim Hendler, and Ora Lassila announced the creation of the Semantic Web, so many of us were very excited to hear Jim Hendler’s update on current developments. Some key themes of his presentation were already covered in this article from August, 2010 in New Scientist: Google, Twitter, and Facebook Build the Semantic Web. With his trademark slogan, “A little semantics goes a long way,” Hendler added some further context, and described how these companies and others have tapped into social and commercial drivers to promote relatively simple approaches to solving the problem of getting content tagged, and thus increasing the ability for computers to understand the meaning of text across vast amounts of Web content.
Read MoreOracle Releases New Features
Oracle has announced several technical and feature enhancements to Oracle Clinical, Oracle Remote Data Capture and Oracle Thesaurus Management System. These upgrades will help clinical trial sponsors and contract research organizations (CROs) launch and conduct global clinical trials more efficiently and effectively.
Read MoreNot True!
The Autonomy folks must be getting worried about the progress of taxonomy applications and the precision and recall that such systems provide. Autonomy and Google live on relevance rankings as the return to the user. Relevance to me is a confidence game. It is the best guess of the system as to whether the results returned will actually match the user’s request. If you have a big enough data set returned, certainly something in there will be useful. But the sheer amount of items the user has to review (or amount of noise they have to look at) is very annoying. So they rank the returns by relevance based on a number of statistical factors so the most likely items based on co-occurrence with terms matches and near matches will appear at the top of the list – that is, they will be relevance ranked.
Read MoreTaxoBank Adds Pasta Lover’s Thesaurus
The first new Taxobank entry in quite some time is the Pasta Lover’s Thesaurus, apparently a student project but nevertheless an informative and mouth-watering example of a straightforward hierarchical and relational thesaurus.
Read MoreWhere Are They Now?
Tech companies come and go. There are always tech savvy entrepreneurs with big ideas looking to fill a need in the market and investors looking to get the huge returns that only come from investing in tech startups.
Read More