Businesses and governmental organizations are increasingly viewing artificial intelligence (AI) as an indispensable technology. To better support the number of AI-based platforms and applications and as usage surges, enterprises have started to deploy dedicated “AI clouds”, optimized for training and inference workloads. Generative AI clouds look set to power innovation across every vertical, from healthcare to retail and public sector services. But what challenges does migration to generative AI cloud entail? and when will generative AI cloud become the mainstream cloud infrastructure for all?
Why is Generative AI different?
While traditional cloud infrastructure was designed for general-purpose computing, storage, networking, web apps, databases and APIs, generative AI cloud infrastructure caters specifically for optimizing, building, training, and deploying generative AI workloads.
According to Fortune Business Insights, last year, 88% of organizations were using AI in at least one business function while 71% were regularly using generative AI. Indeed, investment in AI solutions is expected to soar; IDC predicts that by 2028 global enterprises will invest a staggering $632 billion in the technology.
Given the surge in adoption of generative AI applications – together with the embedding of generative AI to existing applications, known as ‘passive generative AI’, to enhance functionality, demand for specialized cloud infrastructure to facilitate these new services and applications is seeing sharp growth.
Enabling intelligent cloud services
The switch to generative AI cloud solutions is much more than a technical upgrade. By moving to intelligent, generative AI-based tools, many organizations are able to run operations more efficiently and offer new services. Each vertical sector is finding innovative applications of the technology. For example, in healthcare, AI-scribing technology is proving to bring huge benefits to the sector. A major study carried out on behalf of NHS England by GOSH’s Innovation Unit, GOSH DRIVE, found that AI-scribing technology can significantly reduce clinician workload while improving patient care, with potential to unlock millions of pounds worth of activity if rolled out nationally. The results of the study showed an increase in direct patient interaction time during appointments of almost 25%. A&E departments saw particularly strong results, with a 13.4% increase in patients seen per shift.
Similarly, within transport services, generative AI applications are beginning to play a critical role in providing clarity of service delays and security for people using public transport. By gathering live updates on services and monitoring activity on IP-based surveillance cameras that are integrated with the core network, any suspicious behavior (such as a person with a gun on the metro) is detected and reported back to the central AI-based communications hub. From there, the appropriate action can be taken to alert police, remove perpetrators and keep all passengers safe.
The challenges involved in moving to generative AI cloud
A key consideration for IT teams looking to implement new platforms based on generative AI cloud is whether their existing cloud provider can actually support it. Generative AI workloads require fundamentally different underlying hardware and processing architecture compared to traditional cloud computing. At the heart of this shift is the move from CPU-based to GPU-based infrastructure. CPUs process tasks sequentially, which is far too slow for the demands of generative AI at scale. A GPU (Graphics Processing Unit), originally designed to render images and video, is capable of handling massive parallel processing and performing many calculations simultaneously, which makes it well suited to the complex mathematical operations that underpin large language models.
Another very significant concern for enterprises and organizations moving to generative AI cloud is data sovereignty. The conversation has gone beyond simply where data is stored – the real issue is now who ultimately controls the systems that process, analyze, and generate value from that data. For most European organizations, generative AI workloads run on infrastructure owned and operated by US hyperscalers including AWS, Microsoft Azure, and Google Cloud which raises fundamental questions about legal jurisdiction, data access, and operational continuity.
In response to concern from organizations, European authorities are introducing important measures to promote the development sovereign digital infrastructure. The European Chips Act introduced in late 2023 supports semiconductor production, while the Digital Europe program funds advanced digital infrastructure; a €180 million call for tenders is encouraging the development of sovereign cloud solutions to secure the hosting of European institutions.
The future of generative AI cloud
Despite these significant hurdles, there is little doubt that generative AI cloud is on its way to becoming the mainstream cloud infrastructure. Already by 2030 we’ll see new applications coming to market with generative AI built in from day one and within the next ten years generative AI infrastructure will become the norm. The challenges then will be around sustainability and how to reduce overuse of generative AI cloud and keep energy usage to a minimum.