Tuesday, November 5, 2013

Is Risk Really Holding Back Big Data Projects?

In his WSJ article titled "The Risks of Big Data for Companies", Dr. Jordan explains a number of risks that companies should be aware of before jumping into this new and unchartered market.

He outlines that too much data could be dangerous for companies who might invoke inner company politics or make unjustified decisions too early with the results they make. I would imagine there are no risks in ignoring information that the data gives us and just keep moving along as a horse with blinders.

WHAT? If people seriously took the advice of this column in the WSJ, we would never advance in business. More importantly the #1 risk that is keeping up every CEO is not whether a plant in Canada has better output than the same in the US, thusly creating an "inner company rivalry". No, the #1 threat is the competition. The competition can close down your market and make your company's revenue tumble.

Listening to this sort of advice is not only dangerous, but plainly stupid. No company that I have helped with Big Data projects has decided to throw out human intuition and experience for decisions made based on the data alone. Instead all functioning companies want to see whether the data can complement decision making. That is what a pilot will show and the ensuing business case will prove it, or not. Pilot programs should be done with Big Data in small areas to show what we could know, without fear that we will uncover some terrible new insights that could cause inner company conflict.

If you are starting a new Big Data project, please do NOT read this article. Instead follow your own intuition and call me in as an adviser.

Thursday, October 31, 2013

Big Data Projects: Who is the Organization's Big Data Leader

Deloitte Touche Tohmatsu Ltd.
The big question for any consulting organization trying sell Big Data projects is to find the Big Data leader in that organization. There is always one person in every organization who has spurred the company you are selling into to use the massive amounts of data that they have access to to create solid, decision driving data.

Who is that person? Is she the CIO? or is he the CMO? or could she be a Director-level person with a special charter? Tom Davenport, the Data Analytics Guru from Harvard discusses this well within his Wall Street Journal blog. You can read it here. Tom's opinion is pushing for a decision within a corporate organization for where the Big Data Leader should be. I agree with his overall complaint from the Deloitte poll shown below that the leader should be someone who works enterprise-wide, so that an overall Big Data strategy can evolve. In a perfect world that would be great.

But I think reality is different in the companies I have insight to when it comes to their Big Data adoption. In reality, a Big Data project comes out of a need to know more about your customers than your competitors know. It has to start with a dream somewhere in the organization that says, if we had this piece of information, we could sell more, grow market share or save the cost of creating products no one wants. This need could develop in an IT organization, but only if their goals are clearly aligned with the business goals, which sadly is not often the case. Most of the times, I have seen this need arise was in the business unit level, where goals are clearly tied to the success of selling products.

Monday, October 14, 2013

If 64% plan a Big Data project, why are so few IT consulting companies prepared?

-->
According to a Gartner survey, 64% of the companies surveyed plan to invest in some sort of Big Data project in 2013. Many of those surveyed will run some sort of pilot project, not investing in hardware as much as investing time with a singular goal: where will Big Data payoff? how much needs to be invested to get which results?

Despite this strong opportunity, few IT consulting organizations have strengthened their ranks with specialist who can guide their clients through the uncertainty of this new market. One indicator is the lack of career positions posted for Big Data sales specialists and consultants who can create a solid business assessment that is strong enough to justify further funding. There are some who are not only specialized in this area, but also host cloud environments, where their clients can start a real pilot without any data center investments. It seems that these few will do well in the upcoming boom as others scramble for Big Data market share.

Monday, September 30, 2013

Using Video within a Big Data Query for IVS

What is the real value to digital video in a Big Data query? Should video even be a consideration for an other source of data?

Most Big Data applications are out for one simple thing: Getting decision driving information from multiple, disparate, structured and unstructured data sources. If all we had to do is query our RDBMs to get to good decisions, then Big Data as a market would not exist. But it isn't and things like weblogs and twitter feeds - all examples for large sets of unstructured data, are extremely useful for data scientists to use to find decision driving data.

Well, what about video? Can digital video be useful for a query? Early indications are that it is and the more you think about it, the more you find good, solid use cases in the industry. IVS (intelligent video surveillance) is clearly an early adopter to use video. The more that intelligent, digital computer vision algorithms can be combined with queries of other databases to form a real-time, interactive query process, the more intelligible and actionable information we can glean from the query. Yes, there are computer vision algorithms that can search through a digital video file and detect anomalies. But despite the current speed, can we combine that search with a recognition database, like a facial recognition and give back intelligible information. What about height, speed, color, make and model or even an analysis of behavior?

What if an IVS system could automate a query that someone has entered a safe perimeter, that he is recognized as a known person on the terrorist watch list and is spending time cutting through a fence? What security level would this anomaly trip to get action at the highest level. Anyone who claims that any anomaly should trigger highest level does not understand the IVS market. After three or four false alarms, most companies turn off their detection and revert to good, old-fashioned human visual detection.

Big Data has this promise, but video is not an integral part of the query capabilities yet. There is nothing on the market today that allows companies to bring a digital video feed in and query it using SQL like any other data. But that could soon change (stay tuned).

Monday, September 23, 2013

Can we automatically index UGC?

Here is one of the largest dilemas under the current Big Data sun: will be ever be able to automatically index user generated video content (UGC)? Seems like a an impossibility, but would have great benefits if possible.

Many multiple TVs video
Why would we want to create automatic indices and meta data for the uploaded content? Well for one, it would make the mass of user created content that is uploaded to sites like YouTube searchable. Making this mass of video content searchable means we can find the content we are interested in, saving hours of sifting through content that is only saved by the good graces of someone who titled the video correctly or added descriptive meta data to their uploads.

Video is a content rich source of data, since we humans perceive much more data from video than could ever be tagged. But an auto-tagging system could learn over time to get better using facial and behavioral recognition patterns to improve accuracy. This would be an immediate and significant improvement over the voluntary methods we currently use.

The basis for starting automated indexing would be to bring the digital video stream to a standard format where it could be combined with other data sources and included in a standard SQL search. That is now possible, as you can see by my previous posts.

Wednesday, September 18, 2013

Why is transfer of lots-of-data an issue for Big Data?


This may just be a rhetorical question, but there are a few companies developing technologies that allow a massive amount of data to be transported to a DW environment where - using the latest Big Data tools - this data can be analyzed and further used within larger queries. 


One such company is Attunity (www.attunity.com) and their latest product blog post states clearly there solution allows data to boldly go where and when you need it.

What I don't understand is this: Wasn't that the original use for MapReduce that data did not have to reside in a singular DW so that I could query it? As I understand it, this was Google's motivation to build this technology since they clearly understood that all data of the web would not be located on one spot.

According to a paper from TerraData there is a growing issue in the amount of data produced by oil wells and getting that data transferred across the country or countries to a data center is expensive. I guess my simplistic way of design is to ask: Why not put together a couple of servers at the drilling site, build a Hadoop cluster remotely onsite and remotely run queries on it where it is? Isn't it harder to try to store, transport and restore the massive amount of data?

Just a question. Please comment.

Thursday, September 12, 2013

Breakthrough: Hadoop-based Video Analytics

hd_iconPivotal, a company formed from EMC and VMware has created a breakthrough for video analysis. Dr. Victor Fang from their Big Data Science Lab has a demo where he chews through and analyses 5 GB of MPEG-2 video using Pivotal HD and HAWQ in near real-time. His output are exception images, place of movement, speed and many other parameters of the detection.

Why is this big? It will revolutionize security industry's Intelligent Video System market because it will learn to recognize exceptions and feed these to security personnel. Now, the system will intelligently provide information to the monitoring team, rather than hoping to detect with parameter movement algorithms, which have yet to prove themselves.

Watch this video:
http://www.gopivotal.com/resources#http://bitcast-a.v1.sjc1.bitgravity.com/greenplum/pivotalvideo/Unstructured_Data_Video_Analytics_on_Hadoop_Session_1.m4v