Unlike many other big tech platforms, LinkedIn has decided it won’t spend aggressively on expanding its AI data centers this fiscal year. Executives at the professional social network tell WIRED that it plans to keep its investment in GPUs steady, and its compute and storage footprint is also remaining flat. This strategic divergence from the industry-wide surge in AI infrastructure spending signals a new phase in the evolution of large-scale AI deployment, one that prioritizes efficiency and optimization over sheer expansion. The calculations apply to LinkedIn’s fiscal year that began last month and ends next June. The company asserts it has achieved this by finding ways to use its existing GPUs twice as efficiently over the past six months, a feat that could be a bellwether for other organizations navigating the burgeoning costs of AI. While the hardware demands of AI are rapidly evolving, and this plan could still face challenges, LinkedIn executives indicate they have proactively factored in surging prices for essential memory chips.

“One of the goals we’ve set is to try to basically keep our compute footprint flat or as close to flat as possible while shipping more compute-hungry things to production,” says Erran Berger, LinkedIn’s chief technology officer for engineering. “That’s a pretty bold statement to make in today’s world.” Berger and Raghu Hiremagalur, LinkedIn’s chief technology officer for infrastructure, emphasize a commitment to prudent spending, believing these new constraints will act as a powerful catalyst for engineering teams to innovate and develop the numerous generative AI features LinkedIn is poised to launch. Berger further suggests that these efficiency gains are likely to compound, positioning LinkedIn to maximize future data center expansions when budget allocations inevitably increase.

“I really want to double underscore that for a company of our scale, to say a full year we’re going to do this with no incremental storage and compute is no small feat, but it’s taken a ton of work to get there,” Hiremagalur adds, highlighting the immense effort behind this initiative. This approach stands in stark contrast to the prevailing industry trend, where companies like OpenAI, Meta, and Google are reportedly amassing capital and forging unconventional alliances to construct and operate massive data centers equipped with the latest AI-accelerating hardware. The race for computing power has been so intense that labor and parts shortages have hampered numerous projects, leading to limitations on customer access to certain AI tools. Amidst this fervor, growing questions about the long-term sustainability of relentless AI investment are surfacing. LinkedIn, with its user base exceeding 1.3 billion, represents perhaps the most prominent business to publicly address these cost concerns by opting out of the current infrastructure expansion boom.

The Genesis of LinkedIn’s Strategic Shift

The decision to prioritize internal optimization over aggressive hardware acquisition stems from a critical re-evaluation of infrastructure strategy following Microsoft’s acquisition of LinkedIn in 2016. Initially, LinkedIn explored integration with its parent company’s Azure cloud service. However, this proved economically unfeasible for the sheer scale of the professional network. “Microsoft Azure was growing like crazy, the level of customer demand was through the roof, and at the same time we saw skyrocketing growth on the LinkedIn side,” Hiremagalur recalls. This realization prompted a significant strategic pivot.

In 2022, LinkedIn made a decisive move to invest heavily in its own data centers, establishing a robust infrastructure presence in Oregon, Texas, and Virginia. This strategic decision granted LinkedIn unprecedented control over every facet of its technological operations, a move that proved prescient in the face of the emerging AI era. Concurrently, LinkedIn began developing sophisticated AI-powered assistants designed to enhance user experiences in areas such as message composition, job searching, and candidate recruitment. This ambitious undertaking, while promising, was not without its financial implications. “Every query that’s coming to our site has increased in cost over time,” Hiremagalur noted, observing that LinkedIn’s data storage requirements were doubling annually. He candidly admitted, “That is not a sustainable place to be.” This recognition of escalating operational costs served as a primary impetus for the subsequent focus on efficiency.

Mastering Efficiency: The Pillars of LinkedIn’s Optimization Strategy

The core of LinkedIn’s strategy lies in meticulously optimizing data center utilization across the entire AI lifecycle, from the initial training of models to their deployment in real-time user interactions. Hiremagalur’s infrastructure team spearheaded the development of sophisticated measurement tools, providing granular insights into the compute and storage resources consumed by individual development teams. This data enabled the implementation of a dynamic project allocation system within the data centers, ensuring that computing resources remained in active use for longer periods, thereby minimizing idle time. “Our allocation efficiency and utilization of GPUs on the training side is the best that I have seen,” Hiremagalur stated, reporting utilization rates exceeding 95 percent. This level of efficiency is remarkable, especially in the context of AI model training, which is notoriously resource-intensive.

Beyond infrastructure management, LinkedIn has embraced advanced AI techniques such as model distillation. This process involves training smaller, more efficient AI models by leveraging the knowledge embedded within larger, more complex ones. For instance, in its job recommendation tools, a single, streamlined model was developed by learning from two larger predecessor models. This consolidated model not only excels at identifying relevant job openings but also accurately predicts the likelihood of users engaging with those recommendations. The economic advantage of operating these smaller models is significant, yet Berger assures that this efficiency does not come at the cost of performance. “People are finding and discovering jobs that they were not successfully finding before, because the model is doing a really good job of understanding” their professional aspirations, he explained.

Similarly, the AI model responsible for curating user newsfeeds, initially a costly operation, has been significantly optimized. Berger revealed that LinkedIn implemented dozens of improvements, including streamlining model training processes, enhancing the reuse of historical recommendation data, and achieving a more effective balance of workloads between CPUs and GPUs. This holistic approach to optimization has yielded substantial cost savings.

Further demonstrating their commitment to innovation, LinkedIn engineers even undertook the complex task of modifying foundational software on Nvidia processors. These adjustments enabled the hardware to handle tasks exceeding its original design specifications. Additionally, other software components were re-engineered to shift computational burdens from expensive, difficult-to-procure Nvidia GPUs to more readily available CPUs, which also consume less electricity. Collectively, these efficiency initiatives are estimated to have saved LinkedIn approximately $24 million over the past 12 months. This figure is equivalent to powering roughly 1,100 GPUs continuously for an entire year, underscoring the tangible impact of their optimization efforts.

A Prudent Approach in a High-Stakes Landscape

While acknowledging that these savings, substantial as they are, represent a modest fraction of LinkedIn’s $18 billion in annual sales, Hiremagalur stresses the importance of "craft" and the "agility" gained by freeing up resources. This strategic reallocation allows engineers to initiate new projects sooner and integrate more AI capabilities without necessitating an expansion of LinkedIn’s overall computing footprint. Berger further contends that LinkedIn has successfully enhanced its job and candidate matching outcomes while maintaining cost control, thereby generating a positive financial return for Microsoft. “We should be able to deliver better quality by deploying larger models, doing deeper inference, and doing it for cheaper if we can,” he asserted.

Despite its commitment to maintaining flat spending on AI infrastructure, LinkedIn’s data centers are not becoming obsolete. The company has secured commitments for new servers, ensuring a continuous upgrade cycle for aging or malfunctioning equipment. This proactive procurement strategy, coupled with locking in prices ahead of anticipated market increases, has also contributed to cost savings. “The cost of all of this hardware has just gone through the roof,” Hiremagalur observed, noting that some server prices have tripled in recent months, describing the situation as “just nuts.” This foresight demonstrates a sophisticated understanding of supply chain dynamics and market volatility.

Broader Implications: The Dawn of "Tokenomics" and Sustainable AI

LinkedIn’s emphasis on efficiency aligns with a broader industry trend increasingly referred to as "tokenomics"—a more granular analysis of the costs associated with utilizing generative AI tools. Chirag Dekate, who advises businesses on their AI cloud strategies at the consultancy Gartner, observes a significant shift: “Enterprises are evolving from a buy-more era to a do-more era.” He elaborates, “Until now, the mantra was, buy more to save more. But buy more only increases costs.” This sentiment reflects a growing realization that unchecked infrastructure expansion is not a viable long-term strategy for AI adoption.

Dekate notes that even smaller businesses, lacking LinkedIn’s extensive control over their infrastructure, are finding innovative ways to reduce costs. These methods include purging unused software, streamlining workforce, sourcing data center space from specialized "neoclouds" that offer more competitive pricing than traditional providers, and strategically adopting the most cost-effective AI models available for specific projects.

However, Dekate expresses a degree of concern that companies like LinkedIn might encounter limitations if they impose overly stringent financial constraints. The increasing demand for compute and storage in AI development appears to be an inevitable trajectory. “At some point, something has to give,” Dekate warns. “Either you have to compromise on your AI ambitions, or you have to compromise on your mandates to freeze IT spend.” This highlights the inherent tension between ambitious AI development goals and the imperative of fiscal responsibility.

LinkedIn is not entirely abandoning future data center growth; rather, it is fundamentally altering its approach to infrastructure allocation. Hiremagalur indicates that the era of "wild ass guesses" about projected computing resource needs is over. Instead, LinkedIn has chosen "to embrace the chaos" and adopt a more agile, quarter-by-quarter management approach, while remaining keenly focused on long-term return on investment, as Berger explains. This strategic flexibility suggests that while the current period of flattened spending may not endure indefinitely, LinkedIn’s innovative approach to AI infrastructure management is likely to shape its operational philosophy for the foreseeable future, setting a precedent for a more sustainable and efficient AI ecosystem.

By admin

Leave a Reply

Your email address will not be published. Required fields are marked *