AI & ML Automation Data Analytics Governance

Preparing Data for AI

 

As AI reshapes industries, businesses find that preparing data for AI presents challenges that traditional analytics have yet to encounter. AI’s complexity requires a new approach to data management involving careful alignment of architecture, governance, and stakeholder participation. Inadequate preparation can swiftly undermine AI projects, creating risks that leaders must identify and mitigate. Therefore, a solid data preparation strategy is essential to unlock AI’s potential while minimizing its inherent risks.

What specific challenges arise when preparing data for AI? How can organizations ensure they have the tools and frameworks to create a robust, adaptable data architecture? And what ethical and strategic guardrails should be in place to support the evolving role of AI in decision-making?

Contributor

    • Nitin Sharma, Global Director, Partner Business Development, Microsoft
    • Saurabh Jha, SVP & Global Head – Data & Analytics, Tech Mahindra

You might also find these helpful:

Transcript

Sanjog Aul:
Welcome, listeners. This is Sanjog Aul and your host, and the topic for our conversation today is Preparing Data for AI. So what are we talking about here? We have been dealing with AI for many ages for now, but then with the advent of AI, we see the quality and the volume and everything else that has to be looked at, and from a data standpoint, is changing, getting more complex, getting simpler, that is separate, but it’s changing. So what is it that is going to take for us to create a mechanism, an approach, a methodology to prepare data for AI? That’s what we are here to discuss, and what is the outcome? We want it to become as easy as possible.

Sanjog Aul:
We are able to make sure that ethical and privacy concerns are handled properly, and then at the same time, we have the right data architecture. So we are not doing it for today, but for tomorrow. So to discuss this, I have with me Saurabh Jha, who’s the Senior Vice President and Global Head for Data and Analytics with Tech Mahindra. Hey, Saurabh, how are you?

Saurabh Jha :
I’m good, SanjOG. Thank you so much for having me here.

Sanjog Aul:
Glad to have you, and we also have Nitin Sharma, who’s the Global Director, partnerships with Microsoft. Hey, Nitin, how’s life, sir?

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
Exciting. Glad to be here.

Sanjog Aul:
I love the sentiment. Thank you so much for joining her today. So, Saurabh, let’s start with this first question. So why is preparing data such a big deal when we’ve been dealing with data for so many years? Or rather age is what I’ll say, and then, now that we are dealing with AI, how is it becoming any more complex or even more daunting?

Saurabh Jha :
Let me start from the basics. So, for four decades almost, we have been dealing with data the way we have been dealing with it till today, and that is primarily around gathering data from all over the enterprise, bringing it into a central location, what we call as a data warehouse, or call it a big data ecosystem or a data lake and so on and so forth, and then we did some analysis, primarily, which were descriptive in nature, which were static in nature, historical data and so on and so forth. The important point to remember here is who was the audience for this data? The audience or the consumer for this data? Primarily were the people, were the business users who looked at reports, who looked at dashboards and so on and so forth. Then for the past few years now, a new consumer came into the picture, and that is with machine learning models, AI. There was prediction that was required based on historical data, but prediction for the future, and dare I say so, but that model has not matured today. To be very frank, it has not matured to an extent where it can be scaled across the enterprise in most enterprises, even a simple AiML, for example, for an AIML kind of a consumption for predictive modeling, even today, the whole cycle, starting from where you source the data, you store the data, and then you transform it, and then you provide it.

Saurabh Jha :
For AIML, the smoothness or the seamlessness with which it works for Bi and reporting and dashboarding is not the same when it comes to AI and ML, and now we have a third consumer on the obscene, and which is generative AI, which brings all kinds of further complexities into this entire equation. So what I’m basically trying to say is, for the last four decades, we have been handling data for a particular type of a consumer. Last few years, we saw a new type of a consumer, and we have still not matured it, and now we have a third consumer. As you can see, the complexity of managing data is suddenly increasing because it starts all the way from when you source the data. How do you manage it? How do you govern it, how do you ensure the quality? And as all of us are aware, these are challenges which we have been facing for many years, and which we have still not been able to sort out to the satisfaction of everyone. Now, with AI, we have, you know, we have things around ethical, that the AI model should not be ethically biased.

Saurabh Jha :
Then we have things around privacy, that the generative AI model should not be trained with, private data, for example, with PII data and so on and so forth. Also, the AI models have to be explainable. So is our data architecture today that we typically have in organizations. Is it set up in a way to fulfill all of these requirements? The answer is no. That is where the challenges suddenly start compounding from moving from one consumer to three different consumers. Now, which is Bi and analytics. You have AI ML, and now you have generative AI, but the challenges which we have not been able to overcome even for the first use case, now we have two more consumers to take care of.

Saurabh Jha :
So that is why this entire task has suddenly become much more challenging.

Sanjog Aul:
Now, what do you have to say to this?

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
Absolutely, I think I fully agree with what Saurabh said. I would just add that what changes with AI is unlike a classic analytics deployment, which was more of a technology play. All the recent research shows that the biggest blocker in delivering AI projects has been that the company culture is not ready for AI. Why I say that is because we look at AI not just as a technology play, but something that requires full fledged data strategy to be in place, with both the people, process and platform being core tenants to that. If you don’t do that, you want to have difficulty in identifying stakeholders, you’re going to have difficulty in getting their collaboration and participation on your projects. People struggle with identifying what are the business use cases they should prioritize? And there’s the general lack of skills as well as all the legacy data state and quality issues that we inherited from historical data. Challenges become even more compounded when it comes to AI. So it’s a combination of all the data issues as well as the organizational issues that need to be sorted out to make this whole journey fruitful.

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
And that’s really what makes AI so much more daunting when we look at preparing data for AI scenarios.

Sanjog Aul:
So based on your response, Nitin, what do you think should be the first set of steps, or all the steps required in that playbook so that we can start preparing our data effectively and who all need to be involved?

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
Well, you, as I said, you definitely have to have your data strategy at the core of your planning, right? There has to be a plan around how you plan to prepare and manage your data going forward, and bringing your people, process and platform, the three p’s, as we call it, as an integral part of this journey, is super, super critical. So when it comes to people, we recommend that you start with an envisioning workshop that brings all the key stakeholders together. We define their roles and responsibilities and we identify all the Personas, and then the customer can go about figuring out who are the people who will be part of those Personas. You know, typically we see the people coming from four categories. There are the people who are relevant from a data strategy, shared services perspective. So all kinds of architects, you think about enterprise data, cloud security, architects, then you have the data publishers, you have the data stewards, the product owners and architects there.

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
Then you have the data producers where your data science and Mlops leads come in, and then all the data consumers, which is basically all sorts of reporting leaders and data science and ML consumer leaders there, and critical to this journey is also the adoption and change management, which has to be thought through at the ground level because that is important. Otherwise the adoption lags. So once you have the people part of it sorted, then you start thinking about the platform and process part of it. You start by assessing your current state, where really we can help you define what’s the maturity of your data estate today, what sort of governance policies, infrastructure, pre processing and ingestion and integration processes are in place. You start identifying your data sources and how you’re going to integrate them, and then you develop your data strategy with all the nuances that I spoke about. You have to choose your tools and technologies very carefully, because one of the biggest challenges that we hear from all these CDO’s is that they don’t want to be chief integration officers. What’s happening increasingly is there’s such a proliferation of tools and technologies across so many different vendors that most of the time the data teams are spending is in integrating these tools.

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
So you want to avoid that. You want to have strong governance policies in place, and then you also want to take care of drift by having a continuous monitoring and iteration process in place. So all of these things are great ways to get yourself set up to be on that journey of getting your data ready for AI. Saurabh, your thoughts?

Saurabh Jha :
Sure. So I have a little I’ll add on to what you said. So you have already covered all tools, technologies, people, processes. I’ll add a different perspective to this. See, when we talk about data consumption, data, first of all, data production and then data management, and then data consumption generally, what we have observed is in most organizations, when we talk about data, the general thing is that it is a responsibility of the data team, which is a part of the larger IT organization, but if we have to succeed going forward where it is not just about a team of people sitting there getting the data that they are getting, cleansing it, transforming it, and then providing it for reporting here, your data is actually going to determine things which are much more than that. For example, I mean, if you’re going to use generative AI to actually handle a customer situation, you cannot have bad data in there at all. Because if there are, if biases creep in, right, if there are ethical issues, if it’s not responsible enough, it could be a disaster for the organization.

Saurabh Jha :
Now, is it solely a responsibility of the data team to actually take care of this? I wouldn’t fully agree, yes. Data team has its responsibilities to ensure that data is orchestrated properly across all the way from producer to consumer community. However, I think there needs to be a conscious shift in who is responsible for this data, and while from a purely data strategy standpoint, we have seen the movement towards data mesh and domain ownership of data and so on and so forth, but if you ask me, see, data ultimately is created by businesses, the business side of the organization, and it is consumed also by the business side of the organization. So where the data is actually getting created, which is coming to the warehouse lake, which is fine. That’s step number two. The first step is, say, for example, if there are sales folks who are entering data into a CRM system, which is ultimately then finding its way into the AI models, that salesperson needs to be responsible to ensure that the data quality of data that they are inputting is proper. They cannot put dirty data and then expect that there is something great and clean going to come out at the other end of it.

Saurabh Jha :
This is just one example. The same is the case with operations, for example, or with people who are taking care of, who are doing the customer care or customer onboarding. All of that data right at source has to be inputted in the correct manner so that data quality is maintained throughout its lineage and that responsibility has to shift left to the business. So, yes, data mesh, for example, as a part of data strategy, is a step in the right direction, but the entire set of stakeholders, all the way from data creators to data managers, which is the data team and the data consumers, all have to sit on the table and take ownership for their respective parts. So this is, I think, you know, you talked about what needs to change. I think this really needs to change. We cannot just say that, okay, data is the data team’s problem.

Saurabh Jha :
It’s not. Not. Data is business’s problem. So that’s one. The other thing is. So this question is about what are the steps that need to be taken as step number one, two, three. So while Nitin has covered most of it, and I can give examples as well, so there are multiple customers that we work with. In certain set of customers, we again see a very traditional mindset that we have a data team and we have a data science team, which is separate, and the data science team, which basically looks at your AI and Genai and so on.

Saurabh Jha :
When they have a requirement, the data team just hands over the data to them. It’s not a seamless process. It’s a very broken, shoddy handover process that here is the data. Now you play around with it, whereas when it comes to reporting and bi, it’s a very seamless process. Why? Because we have been doing it for 40 years, but the same thing now needs to be repeated for AI and for Genai as well. So from the data team standpoint, when they build the architecture, when they do the tooling, whether it’s for historical data, whether it’s for real time data, and data is getting more and more real time as we go forward.

Saurabh Jha :
And we will see AI and genai being used for a lot of real time use cases, whether it is solving a customer’s problem, whether it is taking decisions in real time, which will have an immediate business impact. So the data team’s mindset there, when they create the data strategy, when they create the data architecture, has to take the seamlessness into consideration. Now, we cannot have seamlessness only where we have Bi and reports concerned. It has to take into consideration AI and Genai as well. So, yeah, these are some of the things which we absolutely have to take care of if we need a future where data AI engineer all work in tandem for the, for business’s benefit.

Sanjog Aul:
So, Saurav, you did lay out a good set of steps, and along with Nitin, you’re kind of giving a playbook, but just because someone creates a playbook doesn’t mean it gets executed to the t or there are no obstacles or traps in the process. So would you be able to share any specific, known and understood traps and nuanced obstacles, obvious and not so obvious obstacles that we could expect in this journey, and then if we do know that these are going to be going to encounter them, what to do about them? How do you get ready for them?

Saurabh Jha :
So, see, when we talk about obstacles, if you ask me, you know, it’s not rocket science. I think we have seen it time and again. People don’t like to change. Right? So the one is ownership. It’s not very, it’s not very easy to imagine business suddenly taking ownership of their data because they have been so used to creating something and then handing it over to the data team that now you manage it, and we want something clean at the end of it. So, if you ask me what I have, what we have generally observed in the last few years is the organization wanting to give that ownership to the business side, but there is a definite resistance in getting in, taking the ownership of that data. That’s number one.

Saurabh Jha :
Number two, the data team itself, like I mentioned in the previous question, when they are designing the data platform for tomorrow, not only for today, they need to have this mindset to ensure that whatever architecture they are putting in place, including the tools, technologies, the workflows, that takes care of all the consumers on the right hand side, not just the traditional set of consumers, which they are very comfortable with. So again, they have to get out of their comfort zone, design their architecture, which takes care of futuristic requirements from data. So, like, you know, in one of your, you just mentioned that it’s as much of a cultural change as it is a technological change and a tooling change, and which I completely agree with. So mindset, the, how do I put it, the ability to take ownership. I think these are the few things, and any organization that needs to succeed in ensuring that data and AI play well together has to undertake a massive change management, change management initiative, and we have seen organizations do it. There are organizations who are very serious about this and they have undertaken massive change management initiative where they have brought everyone together, the business team, the data team, the AI team, and they have formed teams which are a collective of all of these teams and to govern the data and the outcomes all the its generation to its consumption.

Saurabh Jha :
So there are companies which are doing it, but not all of them. So that’s where we are.

Sanjog Aul:
Sanju Nitin, what have you seen?

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
Well, I think Saurabh covered all the data related and organization challenges, related stuff, right? I mean, around data quality, the data silos, governance skills, resources, your data drift and resourcing gaps, overfitting, underfitting, all of that he’s really spoken about. I’ll just touch upon one of the things that we have encountered again as the biggest reason that companies have been slower in adoption, and that really goes back to the foundation of trust. Trust in what they’re doing, trust in what will happen when they adopt this technology, which is intimidating to a lot of people when they’re first thinking about it, and I would say some of the core principles of AI, starting with fairness, reliability. So the system has to be fair to the people in every part. It should be devoid of any biases. It has to be reliable and safe to use.

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
The privacy has to be first class. It needs to be inclusive. The transparency and accountability, these are first class design choices that we need to ensure the customers and business users as well as technical users are comfortable with, and going back to how do we make you ready for these traps? So at Microsoft, we have taken some of these principles and made them standards and implementation processes. So if you think about if you are trying to spin up a new open source model today using Azure AI, one of the first things we recommend you use is the safety service of Azure AI. So all the work that we’ve done to ground our GPT four frontier models in safety is now available to you to start using in your applications, your models and get access to the same level of security and safety and AI readiness that we provide. So it’s really that level of oversight and trust that we’re trying to build, and I’ll also say, in my engagement with customers, one of the concerns that always comes up is what happens to my data.

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
Is it being used to train your models? Is it in any way my information or knowledge getting leaked out? And I would say again, the way you can overcome this resistance or this trap is some of the commitments we’ve made out there, that no customer data is being used to train our models, and all the security around the customer data is always going to be first class enterprise, and on top of it, we also have the copyright commitments we’ve made, that if you use some of our services like OpenAI or Copilot, and you go and do something which potentially exposes you to any sort of copyright infringement, we have provided our commitments on how we’ll secure you around that. So I would say both the data part of the challenges that sort of articulated and the trust part are what you want to envision and bring your stakeholders on board and make them comfortable with the idea that this is not something that’s going to be hard for them to land or get them into any sort of issues, but it’s something that has been well thought out and their design and engineering practices which are available now to take care of these right at source instead of worrying about them when they arise.

Sanjog Aul:
So, Nitin, what do you think we could do to stay ahead of the curve and keep the data ready for AI and not just be playing catch up?

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
Well, great question. I think the first thing to do is to really institutionalize your data strategy, the people, platform and process plays, and then incorporate the learnings from each step. The second, of course, is to drive the responsible AI frameworks that I just spoke about as a design practice in all the AI work being done, grounds up, and then most importantly, I think what we need to think about is how can you actually solve for a use case for the customer and keep repeating it, like make it repeatable? So at Microsoft, we’ve come up with this acronym called Cup Data, which is really solving for a cup of data at a time and then repeating it. So cup data really stands for conforming, unifying, productizing, democratizing, applying, transforming and amplifying. So the C stands for conforming, where you basically talk about bringing the data in through the ingestion, bringing it into a common spot and really kind of processing it to get it to where it needs to be. Then you want to unify all of the data together and then productize it. That’s where you start creating some sort of a data product, which then leads you to democratizing or registering the data as a product, which gives you the ability to apply it to other use cases and other applications on how they can transform it.

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
So you start doing visualizations or analytics or data science on top of it, and then lastly you amplify it by publishing it out using things like purview data catalog, which really helps you to make it available for others to utilize. So it’s a great fun way of kind of looking at the gradual pace of scaling things and then making the process repeatable so you can always stay ahead of the curve and keep the data ready for AI.

Sanjog Aul:
Saurabh, any advice for the world out there about how they can stay ahead?

Saurabh Jha :
So again, I’ll talk in addition tags see, one thing that we have been talking for almost four or five years now is what we call as scaling AI in an enterprise. So enterprise scale AI, or taking AI across the enterprise, but it hasn’t really happened on the ground, and what typically happens is when we talk about AI, or when we have been talking about AI for many last at least half a decade, if not more, is most organizations pick up pocs. So they pick up certain use cases, they create a pocket, and the entire AI strategy basically stays at the POC level, so to say, never scales, and when we go and ask a typical organization, okay, how many, for example, how many machine learning models or ML models do you have in production? It’s typically very limited. The point is, what I’m trying to say is what we need to figure out a way when it comes to staying ahead of the curve is to ensure that we do not always remain at a POC stage, and today with Genai, somehow I think we are kind of at least the organization that we have seen, we are still repeating the same pattern that we are trying to do, pocs, which is fine. See, POC is supposed to be a POC, which means to see once, if it works, if it works, the next big step, which will really give you the competitive edge, is the scale.

Saurabh Jha :
So the only input that I would like to provide here is when you are designing your AI ML strategy or genai strategy, including everything around data, because none of it is possible without data. So the whole thing, from data ingestion to data management, to data strategy, platform architecture, tooling technique, everything to which Genei models you’re going to use, open source, closed source, the whole llmops part of it, or how will you manage or how will you govern your data and keep it ready for, keep it ready for Genai and AI, including responsible data, ethical data, and ensuring, for example, that it should not have bias. For example, when you are first designing it, it has to be designed for scale. It cannot be designed for a POC. A POC is a good thing to start, but it cannot stop there, and that is something we have to avoid. What we did with AIML, it has to, but with a view that it’s going to become so big and it’s going to scale. So your entire.

Saurabh Jha :
And that’s where architecture plays a very important role. The architecture that starts with, you know, data coming in all the way to the place where data is getting consumed, that architecture has to be scalable. If that is not in place. Like I said, you know, we have seen Pos in the last one year. We will see many more pos in the next couple of years, but the scale, the enterprise level scale that we are talking about will elude us.

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
I think it goes back to the earlier question, Sanjok, you were asking about how do you kickstart and get this ready? And that’s the point. You have to have a data strategy which brings together all the key stakeholders, both in the business and the technical side, and makes them part of the journey on what is it that we aspire to do with this? Not just limited to proving the technology, but the people, processes and platform changes which will be required to make this a data driven organization, as we call it, and productize it and make it repeatable and scalable. Because without that, as Saurabh said, you will have point successes, but then the people who really are going to be consumers of that success, they would not have bought into that journey, and it doesn’t then scale as fast as you would have wanted it to.

Sanjog Aul:
Thank you so much again, Saurabh and Nitin for sharing your insights about how organizations can work strategically and on the ground with a mission that they prepare their data for AI effectively.

Saurabh Jha :
Thank you, Sanju. It was a great conversation. Thanks again for having me here.

Sanjog Aul:
Thank you.

Nitin Sharma, Global Director, Partner Business Development, Microsoft:
Thank you so much for setting this up, Sanju, and we definitely would love to discuss more on how we can make this happen and work for organizations together.

Sanjog Aul:
Sanjay Beautiful, thank you so much, and for our audience, please find more conversations like this on CIO Talk Network.

Contributors

Nitin Sharma

Nitin Sharma, Global Director, Partner Business Development, Microsoft

As a Global Director at Microsoft, I lead the partner business development for global system integrators, driving over $1 billion of cloud revenue and consumption impact across Azure, D365, and M365. With more than 23 years of experience in... More   View all posts
Saurabh Jha

Saurabh Jha, SVP & Global Head - Data & Analytics, Tech Mahindra

Saurabh Jha is a Senior Vice President and Global Head of the Data & Analytics (D&A) at Tech Mahindra. This practice helps enterprises Strategize, Design, Implement and Deliver Data & Analytics, Data on Cloud and AI related tran... More   View all posts

Exclusive Sponsors

TECHM - DNA - MPU01 - 300x250
Nitin Sharma