Hot Topics in High-Performance Analytics - InformationWeek

InformationWeek is part of the Informa Tech Division of Informa PLC

This site is operated by a business or businesses owned by Informa PLC and all copyright resides with them.Informa PLC's registered office is 5 Howick Place, London SW1P 1WG. Registered in England and Wales. Number 8860726.

IoT
IoT
Software // Information Management
Commentary
11/17/2008
10:03 AM
Curt Monash
Curt Monash
Commentary
50%
50%

Hot Topics in High-Performance Analytics

For the past few months, I've collected a lot of data points to the effect that high-performance analytics - i.e., beyond straightforward query - is becoming increasingly important. And I've written about some of these topics, including MapReduce, geospatial analytic capabilities and memory-centric analytics among a few others...

For the past few months, I've collected a lot of data points to the effect that high-performance analytics - i.e., beyond straightforward query - is becoming increasingly important. And I've written about some of them at length. For example:

Ack. I can't decide whether "analytics" should be a singular or plural noun. Thoughts?

Another area that's come up which I haven't blogged about so much is data mining in the database. Data mining accounts for a large part of data warehouse use. The traditional way to do data mining is to extract data from the database and dump it into SAS. But there are problems with this scenario, including:

  • There's a lot of data to move.
  • Therefore it's tempting to only sample the database rather than analyze the whole thing, which could have at least a slight negative effect on model accuracy.
  • The result of the process is often some kind of scoring algorithm, and you may want to execute that real-time rather than in batch mode.

Various interesting fixes have been tried.

  • SAS and Teradata are partnering quite closely to run SAS on Teradata boxes.
  • Database management system vendors are building at least the data scoring part right into the DBMS. SAS rival SPSS - which relies more on just-in-time SQL and less on batch extracts anyway - reports that hooking into Oracle's native scoring produces massive performance gains. (To put that another way - I finally got independent confirmation of what Oracle's Charlie Berger has been telling me for years.)
  • Data preparation can be handled by the general ELT/ETLT (Extract/(Transform)/Load/Transform - i.e., in-database data transformation) strategies of the data warehouse DBMS vendors.
  • Oracle (more than most competitors, although SAS/Teradata are headed that way too) actually does all stages of data mining right in the database.

Vendors who are putting considerable marketing emphasis on parallel analytics include:

I'm sure others would say they belong on the list as well. It's an important area of competitive differentiation.For the past few months, I've collected a lot of data points to the effect that high-performance analytics - i.e., beyond straightforward query - is becoming increasingly important. And I've written about some of these topics, including MapReduce, geospatial analytic capabilities and memory-centric analytics among a few others...

We welcome your comments on this topic on our social media channels, or [contact us directly] with questions about the site.
Comment  | 
Print  | 
More Insights
Slideshows
Reflections on Tech in 2019
James M. Connolly, Editorial Director, InformationWeek and Network Computing,  12/9/2019
Slideshows
What Digital Transformation Is (And Isn't)
Cynthia Harvey, Freelance Journalist, InformationWeek,  12/4/2019
Commentary
Watch Out for New Barriers to Faster Software Development
Lisa Morgan, Freelance Writer,  12/3/2019
White Papers
Register for InformationWeek Newsletters
Video
Current Issue
The Cloud Gets Ready for the 20's
This IT Trend Report explores how cloud computing is being shaped for the next phase in its maturation. It will help enterprise IT decision makers and business leaders understand some of the key trends reflected emerging cloud concepts and technologies, and in enterprise cloud usage patterns. Get it today!
Slideshows
Flash Poll