4 Factors That Can Make or Break an AI Project
April 03, 2023

Dmitrii Evstiukhin
Provectus

Share this

Machine Learning (ML) technologies have evolved at an incredible pace over the past few years, and yet multiple studies suggest that most ML projects fail in the real world. Despite the availability of high-quality technologies, there still exist challenges in using these technologies to create and deliver complete solutions, which can be attributed to several factors.

The main causes of failure can be grouped into four categories:

■ failure to frame the ML problem from a business perspective

■ failure to build a team with the right talent, in the right roles

■ failure to select the right data and ML infrastructure

■ failure to properly manage the AI solution in production

Let's dive into each of these areas in more detail.

1. Failure to frame the ML problem from a business perspective

Firstly, failure to frame the ML problem from a business challenge or opportunity perspective is a common issue. Many companies approach ML with unrealistic expectations, or they are simply following the trend to implement ML, without a clear business need or opportunity. This can lead to wasted resources and disappointment when the project fails to deliver the expected results. To avoid this, it is crucial for the ML problem to be clearly defined, with close collaboration between business leaders and experienced engineers. This ensures that both the business and technical aspects of the problem are considered and that the solution is tailored to the specific needs of the company.

2. Failure to build a team with the right talent, in the right roles

The second factor of AI project failure is the failure to put the right talent in the right roles on the team. When a company has a problem to solve, it is important to get the right talent to work on it. However, this can be a challenging task, as it requires the ability to recognize genuine expertise and skill, which in turn requires the presence of that talent within the organization. To address this, companies should invest in training and development programs to develop talent with the necessary skills within the organization. They should also look for external experts who can bring in specialized knowledge.

3. Failure to select the right data and ML infrastructure

The third cause of failure is not having the right data and ML infrastructure. Even with the right talent, a project can still fail if the appropriate data and infrastructure are not in place. Data is the backbone of any ML project, and without quality data, the model cannot deliver accurate results. Infrastructure is also crucial for the success of the project. This includes hardware and software used for data processing, storage, and model training. Without the right infrastructure, the project will be unable to scale and deliver the expected results.

4. Failure to properly manage the AI solution in production

The final major reason for failure is the failure to properly maintain the AI solution in production. This is the final step of any ML project, and it is where many companies stumble. Once the model has been trained and tested, it needs to be integrated into the current business systems, and work at the scale of the business. This requires talent with yet another expert skillset, and it can be challenging to manage the model in production. This includes monitoring the model, updating it as necessary, and addressing any issues that arise.

Essential Capabilities for ML Infrastructure

These four horsemen of AI project failure are common issues that companies face when implementing ML solutions.

The first two issues are not so much technical as organizational. Clearly, when starting such initiatives, the company's leadership should closely watch for any discrepancies in the organizational structure and processes.

The last two factors that often contribute to an ML project’s failure can be attributed to MLOps and can be resolved by an appropriate implementation.

MLOps, or Machine Learning Operations, is a highly fragmented space, and it can be overwhelming to keep up with all the frameworks and platforms available. But there are certain capabilities that are essential for any real-world ML infrastructure solution. One of the most important is scalability. Organizations and use cases often need to be able to scale up and down, to adjust to the usage patterns of end users. Without scalability, an ML solution may be unable to meet the demands of a production environment.

Another important capability is reproducibility. The platform should be able to successfully reproduce an experiment from a month ago, which requires versioning of everything: data, ML code, pipeline configuration, infrastructure code, experiments, and more. This capability ensures that the results are consistent and can be trusted.

Security and observability are also key capabilities for an ML platform. Properly configured security ensures that the data and models are protected from unauthorized access. In its turn, observability ensures that the platform has full visibility into everything, including data, models, infrastructure, code, and users. This allows for a better understanding and management of the solution.

In conclusion, while ML technologies have advanced rapidly in recent years, the implementation of ML solutions in real-world environments remains a challenge. To overcome challenges, companies should clearly define the ML problem through collaboration between business leaders and experienced engineers. They should invest in training and development programs to build the necessary skills within the organization and seek external experts to bring in specialized knowledge.

Additionally, organizations should focus on building a robust ML infrastructure that includes key capabilities, including scalability, reproducibility, security, and observability.

With a well-defined problem, and the right talent, data, and infrastructure in place, companies can increase their chances of success in implementing ML solutions in the real world.

Dmitrii Evstiukhin is Director of Managed Services at Provectus
Share this

The Latest

April 25, 2024

The use of hybrid multicloud models is forecasted to double over the next one to three years as IT decision makers are facing new pressures to modernize IT infrastructures because of drivers like AI, security, and sustainability, according to the Enterprise Cloud Index (ECI) report from Nutanix ...

April 24, 2024

Over the last 20 years Digital Employee Experience has become a necessity for companies committed to digital transformation and improving IT experiences. In fact, by 2025, more than 50% of IT organizations will use digital employee experience to prioritize and measure digital initiative success ...

April 23, 2024

While most companies are now deploying cloud-based technologies, the 2024 Secure Cloud Networking Field Report from Aviatrix found that there is a silent struggle to maximize value from those investments. Many of the challenges organizations have faced over the past several years have evolved, but continue today ...

April 22, 2024

In our latest research, Cisco's The App Attention Index 2023: Beware the Application Generation, 62% of consumers report their expectations for digital experiences are far higher than they were two years ago, and 64% state they are less forgiving of poor digital services than they were just 12 months ago ...

April 19, 2024

In MEAN TIME TO INSIGHT Episode 5, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses the network source of truth ...

April 18, 2024

A vast majority (89%) of organizations have rapidly expanded their technology in the past few years and three quarters (76%) say it's brought with it increased "chaos" that they have to manage, according to Situation Report 2024: Managing Technology Chaos from Software AG ...

April 17, 2024

In 2024 the number one challenge facing IT teams is a lack of skilled workers, and many are turning to automation as an answer, according to IT Trends: 2024 Industry Report ...

April 16, 2024

Organizations are continuing to embrace multicloud environments and cloud-native architectures to enable rapid transformation and deliver secure innovation. However, despite the speed, scale, and agility enabled by these modern cloud ecosystems, organizations are struggling to manage the explosion of data they create, according to The state of observability 2024: Overcoming complexity through AI-driven analytics and automation strategies, a report from Dynatrace ...

April 15, 2024

Organizations recognize the value of observability, but only 10% of them are actually practicing full observability of their applications and infrastructure. This is among the key findings from the recently completed Logz.io 2024 Observability Pulse Survey and Report ...

April 11, 2024

Businesses must adopt a comprehensive Internet Performance Monitoring (IPM) strategy, says Enterprise Management Associates (EMA), a leading IT analyst research firm. This strategy is crucial to bridge the significant observability gap within today's complex IT infrastructures. The recommendation is particularly timely, given that 99% of enterprises are expanding their use of the Internet as a primary connectivity conduit while facing challenges due to the inefficiency of multiple, disjointed monitoring tools, according to Modern Enterprises Must Boost Observability with Internet Performance Monitoring, a new report from EMA and Catchpoint ...