Data Analytics Governance Innovation

Addressing The Big Data Veracity Challenge

Everyone is expecting to get a lot out of Big Data but are we too trusting of the analytics we’re getting out of it? As the variety and number of data sources being harvested increases, there are a lot of steps and many opportunities for errors to creep into the process of creating actionable insights. And when multiple people interpret one set of data very differently, whose interpretation do you run with? What can an organization do to ensure that the insights distilled from the data, can be trusted?
Contributors

    • Philip Wisoff, Chief Information Officer, Proskauer Rose
    • Ram Akella, Professor of IS & Technology Management; Director, Center for Knowledge, Is, &Management of Technology, University of California Santa Cruz

Transcript (Addressing The Big Data Veracity Challenge)

Sanjog Aul [00:00:23]:
Hello and welcome to this segment on CTN, which is CIO Talk Network. To learn more, please visit ciotalknetwork.com and today’s topic is Addressing The Big Data Veracity Challenge and our guests for today’s show are Philip Wisoff, who is the Chief Information Officer with Proskauer. Good morning, Phil, how are you?

Philip Wisoff [00:00:44]:
Good morning. Fine, thank you.

Sanjog Aul [00:00:46]:
Very good. And just before we got started, learned Philip is from New York, so hope things are going well.

Philip Wisoff [00:00:53]:
Yes, we’re recovering, thank you.

Sanjog Aul [00:00:57]:
Very good. And hope everything will settle down in coming days and you can resume your normal life along with your other peers and counterparts and workers in the company. So best wishes there.

Philip Wisoff [00:01:10]:
Thank you very much.

Sanjog Aul [00:01:11]:
And we also have Ram Akella, who is the professor of Information Systems and Technology Management. He’s also the Director of Center For Knowledge Information Systems and Management of Technology with University of California, Santa Cruz. Good morning, Graham, how are you?

Ram Akella [00:01:27]:
Good morning. How are you doing?

Sanjog Aul [00:01:29]:
I am doing fantastic. So how is life at your end?

Ram Akella [00:01:33]:
Wonderful. I’m at Laguna beach enjoying the ocean and the waves before a board meeting and delighted to be working and chatting with you right now and with a nice audience.

Sanjog Aul [00:01:44]:
So now, Phil, can you see how people can tease us with being at such a beautiful and gorgeous location while talking to us about technology?

Philip Wisoff [00:01:53]:
I’m going to recommend that my firm open an office there. Sounds great.

Ram Akella [00:01:59]:
Great, Phil.

Sanjog Aul [00:02:01]:
All right, guys. So the topic that we have picked up, in fact, just to set the premise, we have done other shows on Big Data and the world is talking about Big Data in general, and people are betting their paycheck on what they can get out of it. Now, the question here is that whatever is the end result that we are seeking, that is kind of making a lot of assumptions that you’ll get right clean data, you will have the best processing and timely processing, and then you’ll have the best possible interpretation? There’s so many variables. So when we were to start with this conversation, Phil, I’ll start with you. Do you think can we really bet on an algorithm which basically will pull out some results, and then we will try having the whole business and our enterprise work accordingly? Is it really reasonable to go with that mindset?

Philip Wisoff [00:03:03]:
I think you know the answer. I think ultimately is yes. I’m more concerned about what are you asking the right questions, though? Of the algorithm to get the business solutions that you’re looking for. So I’m very concerned that, you know, the software seems to be somewhat in its early stages. I don’t want to say infancy, but probably in its teenage stages in terms of the analytics but I suspect that that’s going to get stronger. What really to me is the important part is asking the right questions.

Sanjog Aul [00:03:37]:
So Ram, in your world, when you are looking at big data, you’ve dealt with huge amounts of data and you also research and see other organizations playing with it. Where do you think this is today? Is it fully cooked? Is it like we got our act together and the results that we are getting are truly showing us the signs that yes, can be something we can bet our paychecks on?

Ram Akella [00:04:00]:
Well, I think the area is wonderful shape, but it’s got a long ways to go yet. I think algorithms are getting extremely powerful. But I think there are two or three caveats and additional work to be done. First of all, I think as Phil pointed out, what is the right question to ask is important because people think, hey, Big Data, Big Data. But if you don’t know what question to ask, what answers you get, right, so that’s one but even when you have the right question, you have data, you have a lot of noisy data and noisy data can come from two sources.

Ram Akella [00:04:36]:
One is that somehow what is being put in is not right. We’ll come back to that a bit later but you also have lots of data, a lot of which may or may not be pertinent. So how do you sift through all that with a combination of algorithms and humans? That’s a big question that’s being resolved right now and later in the conversation, I’ll give you examples of how we did this at CISCO, when we can take 100 pages of text and reduce it to one page of refined, very useful and pertinent text and so on. So I think it’s very important to process correctly, to ask the right questions and to use humans the most effectively to win out what’s not just right but what’s pertinent to get what you need.

Sanjog Aul [00:05:20]:
Now with that said, isn’t that what you always try to look for? Fill in your world in the legal industry? Because you definitely could if you wanted, you could be flooded with all top types of data. And now with electronic medium you can have structured and unstructured data. So I’m sure this is relevant for you. But do you think if somebody comes up with that earth shattering algorithm or discovery, no pun intended, with the ediscovery process that typically goes on within the law firms, do you think you would still say, okay, I’m going to minimize the human intervention or human interpretation and start looking more and more to an algorithm and I do not see there’s going to be any false positives.

Philip Wisoff [00:06:07]:
Law is a very interesting industry and I think you’ll have just culturally with us, I think it’s going to be a long time before there’s absolutely no human intervention. We may rely fairly heavily, and we already do, particularly with respect to ediscovery, on algorithms to sort through massive amounts of information to find the gems, if you would but even after that’s done, somebody has to look at it. So I don’t think, at least in that area, that we will not have any human intervention and just be totally reliant on algorithms that’s our industry. Maybe in other areas that might be true.

Sanjog Aul [00:06:47]:
Now, Ram, when we look at the very Big Data concept and we are talking about veracity, which is trust, so are we trying to say that we are trusting that it can handle that volume of data or it can come up with appropriate output and then. Or for that matter it can help interpret what exactly is what you want to, which is what we can call as an actionable insight or an action that you can take based on all that crunching, what are we truly betting on? Is it all of the above or one thing in particular?

Ram Akella [00:07:22]:
Well, I think it’s all the above. So in other words, I would say if I’m, I expect an algorithm to do something for me in the field, I’d first ask, is there a way of figuring out whether the data I’m getting is really true or not? And part of it is self consistency, right? Are there self contradictory parts of the data? But if I have a consistent error, then perhaps only a human can see it so that’s the veracity part. The other part is, well, if I’m looking for insights, which is what I really care about, then in terms of insights, I think you certainly want the human to sort of indicate here the sort of thing I’m looking for but the algorithm can also say, hey, in all this data here are some patterns I’m seeing and which of these might be potentially useful for you? Can you let me know? And by the way, the algorithm can also say, look, this part I’m pretty confident about, it looks reasonable to make these sort of statements, but these statements I’m not quite sure and by the way, I need some inputs from you or some additional data to Be pretty clear. So I think these are all the sorts of things the algorithm can be doing in conjunction with a human, expecting some human input as well.

Sanjog Aul [00:08:36]:
Now, Phil, if you were to take Lamborghini, which can run 0 to 100 miles an hour in three seconds, you would want to still and that is presented to you, of course it could be a great toy to have, but then would you look under the hood to see is it safe? Does it really, can that do it consistently? And then will it really give me all that I want besides just speed? So trying to draw a parallel that Big Data is being offered to you and it is like a black box. Would you like to see under the hood or would you just rely on what’s being presented?

Philip Wisoff [00:09:08]:
You know, it’s, you know, it depends. I guess the analogy, if I was to carry it a little further, if I was presented with a Lamborghini in 1910, when autos were just coming out, I may look under the hood but now in 2012, the Lamborghini has a reputation. It’s well understood. I don’t need to look under the hood to do that. I just don’t think we’re there with Big Data and I think Ram is basically saying the same thing. So just give me a black box and saying, trust it is not going to happen now, maybe five years from now, maybe that is what will happen

Philip Wisoff [00:09:49]:
but we are not there yet.

Sanjog Aul [00:09:51]:
So would you say, Phil, that the fact that you’re not ready or basically you would like to look under the hood, does that tell you and others that we should perhaps wait a little bit or you should still start using it, but with the healthy skepticism?

Philip Wisoff [00:10:06]:
I think it’s the latter. I think there’s some real opportunities. Again, I know the law space better than others, but I think there’s some real opportunities out there to take advantage of this. I mean, this is a powerful tool and the mashup of if you would, of kind of structured and unstructured data to glean some business facts that can be useful to running the firm or gaining advantage are great. You can’t ignore it.

Sanjog Aul [00:10:36]:
So Ram, do you think there is a calibration we can do in terms of how much trust or to what degree can we trust the results or the output or the interpretation?

Ram Akella [00:10:48]:
Yes, but that’s both data and domain dependent. You know, earlier on, as alluding to the work of the many companies we are working with, one is a prototype at CISCO which does diagnostics. You know, when people come up with problems with whatever they bought from CISCO and say, hey, is this functioning or not? Can you tell us what’s going on here? We have a problem. How do you know what the problem is, what’s the cause and how do you fix it? Right? So actually our system actually does this mainly automatically or semi automatically with minimal human inputs from the very smart CISCO engineers. So we are validating or verifying how well you’re doing is of the solutions that the system comes up with, are there new solutions that the engineer could not think of by himself or herself? And of the solutions which came up were some missed out that the engineer is providing. So we’ve done those sort of tests and transpires, there are things that the engineer could not have really thought about in the time available or correctly or because the engineer did not know these particular sort of problems because in our algorithms we are fusing knowledge from many, engineers put together. At the same time

Ram Akella [00:11:58]:
there are certain problems that only the engineer could really handle and their algorithm did not bring up.

Sanjog Aul [00:12:04]:
Now, Phil, think about the situation. There are two types of leaders, right? It could be a bean counter and or the investor. So if it’s a bean counter, they would look for efficiency and the CISCO example that RAM just gave, it’s essentially allowing us to reduce the people in support, maybe utilize the details to suggest a solution but you would not break your bank if that solution does not work. So you’re basically reducing the links in that process for a person to get to that end goal and if it doesn’t work, then we just say, okay, let’s try something else.

Sanjog Aul [00:12:40]:
That’s a different type of a situation where Big Data could show promise and it’s a safer bet, perhaps, but how about when people say I’m going to change the way I keep my inventory on a retail shelf based on what Big Data suggests and that reduces the sales it creates as maybe goodwill loss or whatever other major damage. Do you think we should be looking at such applications yet?

Philip Wisoff [00:13:10]:
I think there’s a, I’ll go back to what Ram said. I think there’s different classes of applications. I think for the CISCO example is a great example where by looking at the data, rather than using human intuition to say how did this problem get solved in the past? We have enough data facts to drive an answer to what the problem is and if we can replicate that, then yes, we can do that. The other class of problems though, are the ones that in my sense, we’re not doing something new, but we’re taking what might have traditionally been a human activity and replacing it totally with the data driving the work based on the data, I think that’s going to take longer to do.

Sanjog Aul [00:14:08]:
All right, let’s take a quick break, listeners. We’ll be right back. And Ram, when we come back, we would like to compare this Big Data to business intelligence. We always looked at the data source and if it was not as clean then we would outright reject but in this case we are saying bring it on any type of data that you have, we have got this magic pill and or a factory which will just take everything that you have, inhale and give you those golden nuggets. Isn’t this an apple pie? Let’s explore this when we come back. Please stay tuned.

Sanjog Aul [00:15:43]:
All right, welcome back. So Ram, since you have been dealing with data and its related applications, I’m sure business intelligence is something that is not truly nostalgia it still exists but then that’s how we always looked at it and where we would like to see clean data, we would look at the data integrity, security and very cleanliness of that data and then we would try to say because I have that clean data, we will do certain things on it and that’s what it’s going to get us, reliable results but here we are somehow expecting that no matter what validation may or may not have been performed on the source data, we are going to get good results. Is this really a good approach to how we wanted to basically plan our business and or try to rely on anything and everything just like that, Just because it’s something new, it is a fad factor connected to it.

Ram Akella [00:16:46]:
I think you bring up a great point and I think Big Data is as vulnerable to all sorts of problems as the old BI, especially since we’re trying to do far more than traditional BI because we’re trying to be much more proactive, much more real time, much more action compared to BI. So the point though is if you think of Big Data, I would say Big Data analytics, with the emphasis being analytics is important because part of the analytics that we use is to really look at data coming in because with Big Data, you have several new problems that you didn’t have as much of with business intelligence, BI one of which is sparsity. We always talk of Big Data, right? But just now you spoke of predicting sales and so on. So similarly, whether you predict sales or you look at advertising, you have lots of big data, but you also have a lot of data sparsity. We have a lot of data about how many ads are being seen by users, but very few users actually buy things so there’s a sparsity on the action so unless you clean out the data, which is noisy, and also account and adjust for sparsity, your results may be pretty poor

Ram Akella [00:18:06]:
so there’s a whole component. There are actually four building blocks of doing Big Data analytics, and one of which is something we call extractive analytics and extractive analytics says, given a ton of Big Data, how do you extract what is meaningful before further processing to get insights and action? And so part of that is cleaning, accounting for clean data, or taking dirty or noisy data and extracting the most useful data out of all of it and that’s a huge area of work underway right now and we’re still midway through the process so that’s very important distinction I think you need to bear in mind.

Sanjog Aul [00:18:45]:
So, Phil, in your world, when we look at how the different pieces of the puzzle in a law firm work, and you are custodian of information overall, what is it that the business is looking for or the attorneys in or other people looking for, which would prompt you to say, now this is a compelling case for us to introduce Big Data and we can reliably deliver some results so that we can show ROI on that investment, plus, of course, help the company grow forward.

Philip Wisoff [00:19:17]:
Okay, well, there are two things. There are two areas, at least in the law business where law firms are probably going to be looking to Big Data or already are doing that and there’s the first example is, you know, for the past 50 plus years, attorneys have billed on by the hour because of what happened in the economy in 2008, 2009, plus client demand, plus competition. There’s a lot of price, there’s price competition out there, and lawyers traditionally have not and because it’s been billed by the hour, efficiency was not a chief concern of the lawyers. Now it’s becoming that we have tons of data that we’re generating every day about the hours that our lawyers work and what they’re doing. What we need to do is really pick at that data and pick through that data and understand how to price a particular type of law work, whether it’s a litigation, a contract, or some other type of law work

Philip Wisoff [00:20:19]:
and we believe we can do that and that’s but that’s a, you know, we’re generating all this data to do that. Traditional BI wouldn’t be able to do that for us. The second example is actually a marriage of business intelligence with Big Data and that is we want to look, we want to identify opportunities with existing clients, because that’s where it’s easier to sell on existing clients than capture a new one. There’s lots of information available externally, you know, lawsuits that are set up, various other activities. So if we can marry that and understand where a client may be against a certain legal obstacle, or anticipate that and get to the client, say, we can help you with this, I think there’s going to be a huge opportunity for law firms.

Sanjog Aul [00:21:14]:
Now, where all do you think you are willing to invest today if the environment was right or the funding was available? And where all would you be on the fence? Phil?

Philip Wisoff [00:21:28]:
I think the first example I gave, we’re actually willing to invest in that now, and I think it’s just a question of capturing the data. The second example, I think is a little bit more. This is where you get back to the conversation about algorithms and whether you trust them or not. You know, did the analytics engine spit out a true opportunity with a client? And I think that’s where we’re going to have human intervention. So I think that one’s a little bit further down the road.

Sanjog Aul [00:21:56]:
Ram, if you were to look under the hood across the whole chain, when the data is received, the data is sifted through, the data is processed and or crunched, then the output is interpreted before the final insight. This whole area where all are there gotchas and where all are people using their own respective different flavors, which is making it confusing to see which one is better than the other.

Ram Akella [00:22:23]:
Okay, so let’s just go through step by step, right? So I think that data collection is one of the most problematic areas right now and I think often when people collect data, they don’t know, going back to Phil’s original remarks, they don’t quite know what the question they’re asking is in some areas. So I think identifying what am I trying to answer would be the first point and then collecting the right data that’s one of the biggest gotchas right now. The second part of it is people don’t often so in fact, if you think of litigation and so on, right? And you know, there are several companies working on this, and IBM has what, a billion dollars of revenue from this. So, you know, litigation is also a good example of this. Now, the second part of it is, once I get that how much of human intervention do I need to train the algorithms is a big question so we often find when algorithms and humans have to work together, most people don’t have a clear idea of how much of human inputs are required and what level of expertise and for how long.

Ram Akella [00:23:31]:
So that’s a second gotcha if you’re not careful because people expect this algorithm to be black magic or white magic out of box, and that’s not true, you need to train it. Then the third thing is algorithms often have domains when they’re valid and when they’re not. So, you know, and so if you don’t realize that, for example, maybe an algorithm can only handle up to a million customers, and maybe we have 10 million customers or 100 million customers, so it just comes to a grinding halt. So understanding the scale of an algorithm or when it can work, under what conditions the data is extremely noisy, some algorithms can’t work. So I think understanding each of these can be very important to avoid gotcha situations, I think and often people misuse algorithms, but they don’t quite know what algorithm is relevant, so they just throw algorithms

Ram Akella [00:24:22]:
and that’s not a very correct thing to do and I think we’re still in the process of creating a new generation of data scientists who can figure out what to use when.

Sanjog Aul [00:24:33]:
Now, with all that, you said the first point where you mentioned that we should be careful in collecting data and have the right answers, we should know which answers are to be provided and thus collect data but that doesn’t seem to be the approach. When we are handling the big data concept. They are saying, I’m going to pick up every tweet, every social media post, everything that a customer says, and then without them knowing truly or without soliciting that we’re going to take all of that and somehow try to read between the lines and come up with that magic insight which is going to get us closer to our customer and maybe become even more proactive, before they know what they want, we will know what they want. Is that truly possible? I mean, are we really looking at a crystal ball?

Ram Akella [00:25:16]:
Actually, I think we are halfway there. In my view that it is indeed, true that when you have many sources of information, right, what happens is if you want to understand what a customer or a user wants, right, then there’s what’s called explicit feedback. When you process a lot of data and you actually provide some insights or action to a user and then ask, are you happy or not? And this can be explicit. That’s called explicit feedback. The other possibility is you keep on providing the user either with some results or for example, ads, and you see whether the user reacts favorably or not, very positively and enthusiastically or not. So in that case, the user does not have to say, I like it, I don’t like it. You can just see from the behavior whether the user is happy and enthusiastic, and you use that as feedback and you keep tuning what you keep further providing the user. So with that, actually we can make tremendous headway.

Ram Akella [00:26:15]:
And what we can do is when you look at the tweets, the Facebook inputs, all these different types of inputs, you can process all that and provide what’s relevant to the user. But this is an area of active research right now, and the jury is out yet, because if you see Facebook, the stock not quite tanked, but took a big hit because they couldn’t supply the pertinent ads because they have a lot of data, but it could be a weak signal, or they may not be using the signal very effectively to figure out what the user wants so I think this is midway some people are being extremely successful. As you know, people are a little scared when they get results from Google and others have not been so successful. So we are right in the plum, in the middle of being very successful, but not quite there yet.

Sanjog Aul [00:27:01]:
Phil, in your world, do you think you can read between the lines and even see the emotions involved with an individual who is posting something which you would like to otherwise use in your. In your discovery process or whatever else that you would use that for. Do you think you can truly go to that level in utilizing that big data inference in creating an evidence which could then be admissible and then also is working in your favor?

Philip Wisoff [00:27:35]:
I think it’s a starting. I think all it is is a starting point. As I read through the literature, predicting human behavior is a. Is not. Big data is not great at predicting human behavior. I mean, up to, you know, up to a certain point it can give you an inference, but after that, you’ve got to investigate that and see if it’s really true. So I don’t think we’re going to be able to say, oh, such and such and such and such. A litigation, this person is going to say these things or react this way to this question.

Philip Wisoff [00:28:11]:
I just don’t think we’re going to get that from Big Data.

Sanjog Aul [00:28:14]:
All right, let’s take a quick break listeners. When we come back, let’s look at the very integrity, sanctity security of the source data which is being generated dynamically in real time. So yes, if you were to do this in a batch mode or in an offline mode, if you will, then yes, we can take those careful surveys and then try to get that data inserted. What is it that we could truly do realistically to make sure that all that we are using as a source has its integrity and because all of that is going to also be utilized to validate what the final results are. Please stay tuned. When we are back, we will explore this more.

Sanjog Aul [00:30:01]:
Welcome back. So Phil, I’d like to ask you a question with respect to eDiscovery. That’s a process that is utilized by organizations. They are at all times looking for security and integrity of data because it’s very possible that the output that you come up with the evidence that could be rendered inadmissible in the court of law if somebody challenge or the judge says no, this has got some tampering that has happened. So what is it that in your world, if at all you were to use this big data related analytics and or whatever else that you do in your processes, what would you differently to ensure that you are maintaining the integrity, the security, the sanctity of the very source data that was collected? Is it your responsibility or do you just hope that general counsel and the CIO of the client’s organization is supposed to do that? Because anyway it’s coming to you, right?

Philip Wisoff [00:30:57]:
Well, in our business there’s something called a chain of custody which says the data we have to actually track the data from its source to when it gets on our systems and that process would remain the same. I mean, that’s just a requirement of the legal industry in a broader context. I think when you’re looking at sanctity of data, what I have control over is data that’s generated within my organization and there’s obviously a lot of that but the real issue for me is if I’m going to do some of the other things I talked about when I start pulling data from the outside, how reliable is that data? And then once I have it in, how secure is that data? So those are the issues from my perspective.

Sanjog Aul [00:31:47]:
So would you be willing to hold the bag if something happened at the source, which is with their client? And while there was a chain of custody, could you truly trace it back that I was not at fault, somebody else was, and whosoever was at fault, they’re the ones who are supposed to fix it and or provide you an alternate?

Philip Wisoff [00:32:03]:
Well, that’s the source for a lot of litigations, actually but in all seriousness, the chain of custody comes at this point when we are delivered the data, whether it’s electronically or they physically send us CDs or whatever it is, we log that moment when we have it, and from that point on we have a record of where it is. So that’s just coming into my organization and that is considered the chain of custody. Now, if it’s been tampered with at the originating source, that’s beyond our control.

Sanjog Aul [00:32:37]:
Now, Ram, what do you think can be done in situations where you do not have lot of the same, which is same input in terms of user data is not provided that there’s not very big population, so you cannot build actually a pattern, then do you think Big Data is even useful?

Ram Akella [00:32:57]:
It is. You bring up a very good point that if you talk and advertise, typically these days Big Data is being driven, interestingly by advertising, computational advertising, because you have millions and hundreds of millions and billions of people who are responding and energy analytics. These are some of the areas so in all these cases, you have very many millions of users and you can actually aggregate across all of them to come up with a solution so the individual is only representing the overall population. But if you think of situations when you have an individual, you can do a good job, but it’s certainly much noisier and has more error likely by tracking the individual path of a user and looking at the evolution, then you can do a pretty good job but you still, because you’re at the individual level, you could have a lot more error.

Sanjog Aul [00:33:52]:
So does that really solve somebody’s problem. What could be the use cases or what type of businesses could truly then not use Big Data because of this very shortcoming?

Ram Akella [00:34:03]:
Well, you know, it’s interesting. So if you think of and I would frame it as not being able to use Big Data in an easy way, but needing much more hard work, I think Phil’s example of legal is a very good example. If you want to have individual think of it. If you have an automated home and you walk into the house and you want the temperature at the right setting, the lighting at the right setting, all these things, I think some components of that you can use big data to see who are other similar users, but the rest of it, you have very idiosyncratic behavior is something which will need a lot of hard work. Anything, for example, any creative activity of creating a new product when there’s something which you cannot quite predict is much harder. I mean, you can enable. But if every individual has a different style of doing things, something creative, it’s much harder to use big data.

Sanjog Aul [00:34:58]:
Now we always look at the human side of any technology implementation. In this case more than anything else, we are still relying on the final interpretation of the output that is created through that big data churning machine. And those are the people who could interpret it differently. So I give the same data set to 20 people, maybe with similar or say a different background and or experience level, they may come up with more or less or different insights. So Phil, do you think if that is what it is, then what are we getting out of this and in your space? So If I had 20 different attorneys looking at the same case with some data presented to them, if they are to come up with different interpretations or different, different case models, is that really leaving technology redundant? Almost.

Philip Wisoff [00:35:52]:
Well, I think the potential, and we haven’t gotten there yet, particularly with what you’re talking about, is for the analytics to make us more efficient with respect to reviewing the information. And we’re already starting to do some of that in terms of when we get a discovery case and we could get terabytes of data, millions of emails, millions of documents, videos, audio clips, we need some way to manage all that information and again, get down to the pertinent pieces. And that’s really the starting point. Once we said, okay, out of this huge set of data, we have a thousand items we should look at rather than millions of items to look at. That’s where we’re getting the efficiencies.

Sanjog Aul [00:36:49]:
Now do you think RAM is the people side, the weakest link in Making big data successful. Because at the end of the day, these are the people, number one, you would not know suddenly how many people of reliable skill levels are available to do something like this. And then you don’t have a benchmark today for you to say this person is really delivering good insights versus not as good insights. Do we need to fail a little more for us to figure out who is good versus who is not?

Ram Akella [00:37:21]:
That’s a great question. I love your question. Because you know, Sanjog, this term crowdsourcing is now increasingly being used in connection with Big Data analytics. So in other words, when you have lots of people who are actually providing inputs to enable algorithms or to train algorithms, how do you make it work? And so far, most people, the assumptions being somehow these people are all interchangeable, which is not the truth at all and they actually come in varying degrees of either levels of expertise in the same area or expertise in different aspects of assets and so I think the right choice of these people, depending on the accuracy level you need. So let’s say you might want better and better accuracy in terms of insights, right? Well, it depends on what’s my economic value. So if I’m going to get a lot of economic value by processing the data and coming to insights

Ram Akella [00:38:17]:
along a certain dimension, then you certainly want to have smarter people with more experience for helping me whom I may want to pay more. Whereas if something is not as valuable, then I may not want to kill myself either on the algorithm front or the person front. Right. So that is very much underway right now.

Sanjog Aul [00:38:35]:
So, Phil, coming to you, do you actually mentioned in your previous comment that this could be utilized to make you efficient? How about being effective? Do you think this particular whole process or technology, or this combination of people, process and technology has the potential to make an organization effective versus just efficient?

Philip Wisoff [00:39:02]:
Absolutely and I would go back to the earlier examples of where we’re really looking for big data to help us either doing a much better or more effective job of pricing our work and identifying opportunities and helping our clients. So I think yes, the answer is definitely yes. So it’s not just an efficiency thing.

Sanjog Aul [00:39:31]:
All right, so then in your world, when you are trying to make people effective, if at all you try to do that, what all would you like to. What all is left to be desired from Big Data? What would you see where people say, yes, this is something which is not helping me save a buck? It’s going beyond that.

Philip Wisoff [00:39:53]:
Like, I think this goes back to, you know, can you predict what individuals are going to do? I Just don’t think that’s where this is going to be viewed as having much value. I think if it helps us run the business better and more effectively or present efficiencies, that’s where I think the value. That’s where the value is in terms of my perspective.

Ram Akella [00:40:17]:
Hello?

Sanjog Aul [00:40:18]:
Yes, hi, Ram. Yeah, we’re finishing up with Phil and definitely we’ll be having you continue your answer. Yeah, go ahead, Phil.

Philip Wisoff [00:40:27]:
No, no, I finished.

Sanjog Aul [00:40:28]:
Yeah, okay, good. Now, yes, so Ram please share your views. I think we lost you when you were about to share a particular angle to the answer that you were about to provide.

Ram Akella [00:40:40]:
Okay, I was just saying that I don’t know whether you heard the comment I made about crowdsourcing of humans.

Sanjog Aul [00:40:49]:
Yes, yes, you just started with that.

Ram Akella [00:40:52]:
So the thing is, you have humans with different levels of expertise and they cost different amounts of money and they’re very helpful in actually training and guiding algorithms but it then depends on what is the value. If I get a better answer, am I going to make much more money or benefit much more? And then it’s worth having an expert. So I think characterizing experts, either by the depth of their expertise or the different facet of their expertise makes a huge difference in training these algorithms. So the human becomes very important and I think evaluating what they can do to enhance the answer makes a big difference. Now this also relates, by the way, Sanjog, to question you asked earlier, because humans are very good at disambiguating when there is something that’s not very clear and where the algorithm doesn’t quite know what to do. Also, I think when you predict, you base it on data, but not only is there the current data, but a human may actually have additional expertise based on a lot of data exposure in other contexts

Ram Akella [00:42:00]:
and that is what the human is bringing in compared to the algorithm and that’s why the human is so important.

Sanjog Aul [00:42:07]:
All right, great. Now let’s take a quick break. When we come back, Phil, I’d like to ask you about the influence management that an IT leader might have to demonstrate because if you were to implement Big Data, yes, miracles are being expected, or maybe some output is expected. But then the people who are creating data or are responsible or being custodians of data, and then the ones who are churning the results and then the people who are interpreting, there are two or three different camps and all of them have to coordinate, they have to work together and then go back round in circles couple of times or iterations are to be done before a final good result can be produced and used. You may be at the top and trying to manage it all. What would it take for you to do that? Please stay tuned. We’ll be right back and explore.

Sanjog Aul [00:43:57]:
Welcome back. So we’re talking about influence management. So Phil, if Big Data is such a process which is touching business, is touching operations, it’s touching it and you might be seen as someone who is running because of being an IT leader and you can at most influence some of the entities and or people, how do you expect to pull this off and do it on an ongoing basis successfully?

Philip Wisoff [00:44:22]:
Yeah, I think I have. I think it’s my as a CIO, I think it’s my responsibility to present the possible to the appropriate management people. I think Big Data to me in terms of the organization is really a mashup of the business development group, the finance group and IT working together because I think the kinds of business problems we’re trying to solve is going to take expertise from all those groups, frankly, and I’m not sure where this ends up. I think over time we’re going to see if, the promise of big data evolves the way I think it’s going to. I think you’re going to see almost a separate entity spring up in most organizations that deal with this because it’s not just an IT discipline.

Sanjog Aul [00:45:14]:
So Ram, in your case, when you do see organizations successfully or not as successfully delivering on the big data initiatives where all from a leadership and management and oversight and governance standpoint, things fail.

Ram Akella [00:45:30]:
I think first of all just task assignment. I think identifying the person who drives the initiative. I think people often get confused between the infrastructure things like Hadoop with the analytic capability. So I think choice of the leader actually and the person who actually understands the totality is one place where things often tend to fall apart. The governance issue. I’ve always found organizations, if I have a CIO who could potentially become a COO and a CEO, that is that business people, not just technologists, makes a big difference. Whereas if they report into CFO and people view this as more like a cost center or something to be managed, then also things may or may not work.

Sanjog Aul [00:46:24]:
All right, so Phil, in your world, do you have any message for the community out there which is trying to get miracles out of this process? I do not see Big Data as a technology alone because we are looking at a complete phenomena. So what is your message to those people?

Philip Wisoff [00:46:43]:
Well, I think there’s some excellent literature out there. If somebody is trying to learn about this, I think there’s some excellent literature out on the net that you can pick up. One article in particular kind of gives you a good overview on how to, how organizations might take advantage of this is there’s a Harvard Business Review article that came out recently, I think in October that kind of, it’s fairly easy to read and kind of sums it up for on the business side. I think that, you know, the other thing is you just need to get yourself educated about what it is and then think about the possibilities with respect to your company or organization and finally, there’s literally thousands of use cases out there that have been documented or discussed so that you don’t have to reinvent the wheel. So there’s, lots of information out there for you to get educated about.

Sanjog Aul [00:47:43]:
Anything, which you would like to caution people who are trying to take this on so that they don’t burn themselves or they don’t get overzealous.

Philip Wisoff [00:47:54]:
I would talk about the hype cycle. Like with any new technology or thing that’s very popular, it’s not going to be the answer to all questions.

Sanjog Aul [00:48:06]:
All right, now, Ram, one final question for you. So as you see organizations out there trying this, where do you see the future going and how far are we from actually establishing the big data veracity?

Ram Akella [00:48:25]:
This for me. Okay.

Sanjog Aul [00:48:26]:
Yes.

Ram Akella [00:48:27]:
Yeah, I think, there’s tremendous potential. I should actually say there’s so much potential as described in the McKinsey report and also in the HBR article by Eric that Phil was referring to, that I think people are hiring data scientists in huge volumes and we’re actually starting new programs at the University of California, both at Santa Cruz and Berkeley in this area and the emphasis there is really asking the right question and in terms of the veracity part of really establishing processes and procedures so that given questions, what are the data you require and what sort of testing might you do both data analytically and in an organizational sense to make certain that the data is useful and productive is something that is underway right now.

Sanjog Aul [00:49:19]:
On behalf of the show and our listeners, I’d like to thank you, Phil and Ram, for sharing your thoughts about addressing the Big Data Veracity Challenge.

Philip Wisoff [00:49:28]:
It was a pleasure. Thank you.

Ram Akella [00:49:30]:
Real pleasure. Sanjog and Phil, thank you all.

Sanjog Aul [00:49:33]:
Thank you and listeners pease like us on Facebook, search for CIO Talk Network and be sure to follow us on Twitter and LinkedIn. Thank you again for listening to this segment on CIO Talk Network. This is Sanjog all, your talk show host till next week. Take care and God bless.

Speaker A [00:49:53]:
Thank you for tuning in to CTN CIO Talk Network with your host Sanjog Aul. To learn more about our program or for show archives, comments or questions, please visit ciotalknetwork.

Speaker A [00:50:07]:
Thank you again for listening.

Download Podcast

Google Podcast, Spotify, Pandora, iHeartRadio, SoundCloud, TuneIn, and Stitcher. Find other syndication channels here or search CIO Talk Network podcast on any other app.

Explore More

Contributors

Philip Wisoff

Philip Wisoff, Chief Information Officer, Proskauer Rose

Phil Wisoff, Chief Information Officer of Proskauer, has more than 25 years of senior management experience in information technology in industries that include law firm, accounting, insurance, telecommunications and pharmaceuticals. Phil f... More   View all posts
Ram Akella

Ram Akella, rofessor of IS & Technology Management; Director, Center for Knowledge, Is, &Management of Technology, University of California Santa Cruz

Professor Ram Akella is currently Professor and Director of the Center for Large-scale Live Analytics and Smart Services (CLASS), which includes SMART (Social Media Analytics Research Transformation), and the Center for Knowledge, Informati... More   View all posts
Add Comment
Click here to post a comment

Advertisement

Persistent - HiTech- MPU - 300x250
Philip Wisoff