Every technological revolution produces a construction boom. Railways. Fiber. Server farms. The boom always begins with a real shortage, then turns that shortage into a story about permanent demand.

AI has reached that stage. Governments compete for data-center investment. Utilities prepare for a surge in electricity use. Investors treat access to power, land and GPUs as the new industrial advantage.

They have evidence on their side. The International Energy Agency expects global data-center electricity consumption to rise from roughly 485 terawatt-hours in 2025 to around 950 TWh in 2030. AI-focused facilities are growing faster still.

But forecasts like these describe a trajectory. They do not guarantee its shape.

The current boom rests on a quiet assumption: that the place where we run AI tomorrow will look like the place where we train it today. I think that assumption is wrong because it ignores one of the strongest instincts in business: control the assets that matter most.

Training is not inference

Training a frontier model is an industrial process. It requires vast clusters of accelerators, sophisticated networking, enormous power supplies and teams capable of keeping the whole machine working. That will remain centralized.

Most businesses are not training the next frontier model.

They are searching contracts, summarizing meetings, reviewing source code, retrieving internal knowledge and running agents against company data. Those jobs are inference. They happen repeatedly, often involve sensitive information and do not automatically require the most capable model on Earth.

Today, the cloud is usually the easiest place to run them. Ease has won while local alternatives remain expensive, fragmented or less capable. That advantage will not survive unchanged once a workstation can perform the same useful work at an acceptable price.

The clue is memory

AI hardware is marketed through compute: more cores, more operations, more impressive numbers. Inference often has another constraint. A model's weights and working context have to fit somewhere, and the system has to move that data fast enough to generate each next token.

Research into large-model serving shows that memory bandwidth can remain the bottleneck even when expensive GPU compute is underused. The exact balance varies by model and workload, but the lesson matters: a machine does not need to beat a data-center GPU at every operation to become useful. It needs enough accessible memory, enough bandwidth and acceptable speed for the job in front of it.

Apple has already demonstrated a 670-billion-parameter model running locally on an M3 Ultra Mac with 512 GB of unified memory. That is a technical demonstration, not a normal office deployment. Still, it establishes the direction.

Now comes an even more provocative signal. Reporting attributed to Bloomberg's Mark Gurman says Apple is designing an M7 Ultra that could support as much as 1.5 TB of unified memory. Apple has confirmed neither the chip nor that configuration, and the report says availability would depend on the memory market. Treat it as a rumor, not a product announcement.

But pay attention to what the rumor describes. The headline feature is not simply a faster desktop. It is a desktop-class system designed to hold extraordinary AI workloads close to the user.

Companies will pay to own the answer

Imagine a law firm, hospital, bank or game studio buying one or two high-memory machines for its internal models.

Contracts, patient records, customer files, strategy documents and source code do not have to be shipped to a distant data center for every request. Common jobs can run inside the organization's own security boundary. Capacity becomes predictable. Network dependence falls. Heavy, steady usage can be paid for through owned hardware rather than a metered API.

The decisive calculation will not be token price alone. Corporate secrets have value. Regulatory exposure has value. The ability to decide exactly where data is stored, which model can inspect it and whether any external provider can retain it has value.

Local systems still need maintenance and security. Companies already accept those responsibilities for identities, laptops, networks and critical databases because surrendering control is not always the cheaper option. AI will enter the same calculation.

The question is not whether the cloud disappears. It is how much valuable work never needs to travel there.

Apple is a clue, not the market

Apple's own architecture follows this logic: process requests on the device where possible, then use Private Cloud Compute for work that needs a larger server-side model. Its high-memory hardware points toward machines that can keep increasingly serious models close to their users.

But Apple will not be allowed to own this category alone. PC manufacturers, chip designers and an open-source community are assembling the same future from the other direction. Open-weight models can already run through tools such as Ollama, llama.cpp and vLLM. Self-hosted platforms combine files, photos, calendars, automation, search and AI agents into personal clouds running on hardware their owners control.

PewDiePie is an unexpectedly useful example. In 2025, he showed millions of viewers how he was replacing Google services with self-hosted alternatives. He then built a local multi-GPU AI system, experimented with open models and coding tools, and released Odysseus, an open-source self-hosted AI workspace with model downloads, shell access, research and agent features.

One celebrity's home lab is not an enterprise trend report. It is something more interesting for a prediction: a glimpse of previously specialist technology becoming a culture. The important part is not his particular hardware. It is that storage, applications, models and agents are beginning to form a coherent private ecosystem in an ordinary person's home.

What enthusiasts build in closets today often returns as a product companies can buy tomorrow.

Where the bubble could form

Data centers are financed over long horizons. The danger is not that AI demand vanishes. It is that developers build for the wrong mix of demand.

Total AI use will explode. Models will also become more efficient, capable open-weight systems will fit on workstations and small office clusters, and privacy requirements will push companies toward equipment they control. Centralized training can keep growing while a meaningful share of routine inference migrates outward.

That means AI can consume much more compute while some hyperscale capacity built for future inference still struggles to earn the return its investors expected.

That is my prediction: by 2036, a portion of the AI data centers planned during this boom will look oversized—not because AI failed, but because inference moved closer to the people and data using it.

The market currently prices one future with extraordinary confidence: more AI means more centralized infrastructure. Technology history rarely rewards that kind of straight-line thinking.

The companies that win the next phase of AI will not only build enormous data centers. They will sell the machines, chips, models and software that let customers avoid them. Some of the most consequential AI hardware will sit under a desk, in a server closet or in the next room—keeping corporate knowledge where its owners believe it belongs.