Software
Commentary
5/13/2008
12:04 PM
Connect Directly
Twitter
RSS
E-Mail
50%
50%
Repost This

The Enterprise Future Of Semantic Search

Powerset launched a tool to search Wikipedia and open source database Freebase Monday, but the technology that powers the search startup could wind up at home in a corporate setting.

Powerset launched a tool to search Wikipedia and open source database Freebase Monday, but the technology that powers the search startup could wind up at home in a corporate setting.Powerset specializes in what's become known as natural language or semantic search. Rather than relying exclusively or primarily only on linking algorithms, as Google and other major Web search engines do, Powerset also uses an "ontology" of syntax, grammar, and sentence structure and, to a lesser extent, thesauri, in an attempt to pull meaning from queries and Web pages -- or in the case of businesses, eventually documents and files. "Search is only a small part of the engine here," Powerset CTO Barney Pell said in an interview. "Really, it's a content understanding engine."

Other search companies are getting in on the semantic game as well, and are looking toward businesses as potential customers.

• Startup Hakia has begun targeting users searching for legal, medical, and financial information, and licensed its technology to a startup that summarizes information for law enforcement, government agencies, and pharmaceutical companies. • Semantra is a new enterprise search start-up focused exclusively on semantic search, and does what it calls "conversational analytics" for Microsoft CRM and all major relational databases. • Q-Go is a Dutch company that does natural language in several verticals. Its customers include DHL, KLM, and Deutsche Telekom. • Cognition similarly does vertical search with a natural language angle, and has its own Wikipedia search. • Inquira uses natural language search in its customer service app to help support staff answer broad or unclear questions, and counts among its customers Honda, SunTrust, and Honeywell. • Astute Solutions' RealDialog uses natural language processing for Web self service support • Even the major Web and corporate search companies, including Google, IBM, and Microsoft, have their own semantic search efforts under way.

Powerset can distill a Wikipedia article into key concepts by picking out verb-noun relationships. A search for David Lee Roth, for example, brings back the information that he was the lead singer for Van Halen, later left the band, and at some point released an album called Skyscraper. Since the engine also uses Freebase -- it could do the same with, say, a drug catalog or a customer database if they were structured properly -- the results page also returns a short bio for Roth.

Since semantic search engines can pick at meaning, they also help when the searcher has a broad question that couldn't be answered if they didn't understand concepts or words. The classic Powerset example is a search for "politicians killed by disease," which brings back the information that, for example, Benjamin Harrison died of a combination of the flu and pneumonia. Imagine querying a business search engine for something like "customers in a hurricane zone" or "employees in senior management."

There are many business scenarios where information is highly specialized, rare, and often placed into odd structures. There's a ton of specialized data in businesses, but Powerset doesn't have to know the meaning of a word to understand related concepts, so it and search engines like it could be of use in the corporate world. Semantra, for example, might return a CRM query for "which retail accounts in Baltimore have sales opportunities of more than $50,000 to women before next week?" or something like it.

There's no dictionary defining an iPhone for Powerset, or that it is made by Apple, or even that it's a device, yet a search distills facts about it, showing, for example that it supports Bluetooth, a fact that's pulled from the Apple Wikipedia entry, not the iPhone's entry. Additionally, semantic technologies can also help return results when information is rare or poorly labeled, because it doesn't rely on links.

That's all the good news. The bad news: Powerset's far from ready for prime time, and neither are some of the other semantic search engines. Google still beats Powerset hands down on a number of queries, and other semantic search engines only have limited rule bases or dictionaries backing them up. Try before you buy is the operative here.

Comment  | 
Print  | 
More Insights
Google in the Enterprise Survey
Google in the Enterprise Survey
There's no doubt Google has made headway into businesses: Just 28 percent discourage or ban use of its productivity ­products, and 69 percent cite Google Apps' good or excellent ­mobility. But progress could still stall: 59 percent of nonusers ­distrust the security of Google's cloud. Its data privacy is an open question, and 37 percent worry about integration.
Register for InformationWeek Newsletters
White Papers
Current Issue
InformationWeek Elite 100 - 2014
Our InformationWeek Elite 100 issue -- our 26th ranking of technology innovators -- shines a spotlight on businesses that are succeeding because of their digital strategies. We take a close at look at the top five companies in this year's ranking and the eight winners of our Business Innovation awards, and offer 20 great ideas that you can use in your company. We also provide a ranked list of our Elite 100 innovators.
Video
Slideshows
Twitter Feed
Audio Interviews
Archived Audio Interviews
GE is a leader in combining connected devices and advanced analytics in pursuit of practical goals like less downtime, lower operating costs, and higher throughput. At GIO Power & Water, CIO Jim Fowler is part of the team exploring how to apply these techniques to some of the world's essential infrastructure, from power plants to water treatment systems. Join us, and bring your questions, as we talk about what's ahead.