In our last deep dive, “AI Workloads at the Telco Edge,” we looked at AI inference as a workload-placement problem. Rather than treating “the edge” as a single place, it broke deployment into infrastructure tiers, from regional data centres and metro nodes to enterprise sites and far-edge environments, each with different constraints around latency, bandwidth, power, cooling, and cost.
This podcast takes the next step. Instead of asking whether AI should move to the edge, it asks a more useful question:
which workloads genuinely need to run close to where data is created, and which are better handled in larger, more economical data centres?
This episode is part of my nine-part series, Build vs Buy vs Rent: The AI Infrastructure Decision Tree for Startups. I recorded it a few weeks ago, before publishing my recent deep dive into AI inference infrastructure, but it connects directly to that work by examining the practical reasons organisations choose to run inference locally in the first place.
The compliance and network reality and - why data residency, security, connectivity, and operational control can drive edge adoption.
Why organisations choose to run AI closer to the data
Two forces push AI workloads toward the edge, and they’re often conflated.
The first is speed:
when a system is controlling equipment, supporting a safety-critical decision, monitoring a live process, or reacting to a real-time physical event, waiting on a round trip to a distant cloud simply isn’t practical.
The second, less discussed, force is constraint.
Data often has to stay local because of privacy, data residency, security, commercial sensitivity, or unreliable connectivity. A factory, hospital, transport operator, utility provider, or critical-infrastructure site may not be comfortable — or even legally permitted, to send every video stream, sensor reading, or operational record to a central cloud.
Together, these forces reframe the edge-AI question.
It stops being about chasing the lowest possible latency and becomes about balancing three competing needs:
keeping sensitive or high-volume data close to where it’s created,
running models reliably where they’re needed, and
managing the real cost, power, and operational burden of local compute hardware.
Get this balance right, and organisations gain faster decisions, lower bandwidth use, more resilience during connectivity problems, and stronger control over sensitive information, without the assumption that everything must move to the cloud by default.
None of this is free, though.
AI hardware needs power, cooling, physical space, maintenance, and careful monitoring, and the more capable the model, the greater those demands become. That’s why the real decision isn’t “cloud versus edge” - it’s deciding what must stay local, what can run nearby, and what can still be handled centrally.
The business case for moving inference to the edge
For founders building AI for physical environments such as manufacturing, transport, energy, agriculture, and retail, a cloud-first default can become the wrong architecture when workloads are persistent, latency-sensitive, connectivity-constrained, or subject to local data requirements.
for example lets consider an illustrative deployment with 1,000 inferences per hour across 500 locations.
A cloud-based approach can become expensive when large volumes of sensor, image, video, or event data must be processed continuously across hundreds of locations. By contrast, an edge deployment shifts more of that cost into upfront hardware, local power, device management, maintenance, and lifecycle support.
The break-even point varies significantly by model size, inference frequency, batching, bandwidth, hardware choice, energy pricing, support requirements, and expected device lifetime. But for sustained, high-volume workloads, the total cost of ownership can favor edge or hybrid architectures over routing every inference through a cloud service.
The exact numbers depend on the model, inference volume, cloud pricing, hardware lifetime, local power costs, maintenance, and utilisation, but the underlying shift is what matters: recurring API spend becomes owned infrastructure and operating cost. For high-volume, latency-sensitive, or data-sovereign workloads, edge and hybrid architectures can materially reduce five-year costs compared with routing every inference to the cloud.
A useful test is to ask whether at least two of these are true:
the application needs a consistently low-latency response; each location generates sustained volumes of data or inference requests; data residency, privacy, security, or commercial requirements make cloud routing difficult; or the application must keep running when connectivity is unreliable.
New local hardware is also expanding what’s possible outside a central data centre, systems such as NVIDIA’s Jetson range and desk-side AI development systems point toward increasingly capable models running in compact local environments. . But hardware capacity alone doesn’t make an edge deployment viable;
power, cooling, model optimisation, network resilience, monitoring, and maintenance still matter just as much.
Listen to the full episode here, on Substack app, or Apple, Spotify / youtube.
In the coming weeks, I’ll publish few deep dives on the decisions, constraints, and trade-offs that shape practical edge-AI deployments, so this episode and previous AI workload deep dive become part of a wider series on AI inference infrastructure covering it in 360 degrees- from where workloads should run, to the practical realities of deploying and scaling them outside centralized cloud environments, so subscribe to the newsletter.
















