In part one of this series, you read why AI is breaking assumptions that have held for two decades. Traffic is centralized, actors on the network are mostly human, and the network’s job stops at moving bytes reliably.
If you’ve not read that article, the TL;DR version: AI has opened a visibility, governance, and operational challenge that today’s network architectures were never designed to close. Now we can talk about the approach to addressing it, a different way of architecting the network for a world where most of the traffic is generated by software, in addition to people.
The Shape of the Network Is Changing
- Uplink and downlink are converging. A chat prompt, an agent’s tool call, and a retrieval-augmented lookup all push meaningful payloads both upstream and down. Traffic that was download-heavy is more symmetric, and east-west (model to model, agent to tool, service to service) rather than north-south flow between a user and a data center.
Compute, and therefore traffic, is distributed by design. Inference runs across hyperscalers, specialized neoclouds, on-premises clusters, and edge locations. What used to be a single, centralized path to an app is now a fragmented mesh of paths to the right model, wherever it happens to be running today.
The traffic pattern has changed shape. The latency envelope that WANs were tuned for has effectively collapsed. Static transactions have given way to streaming sessions over TLS/websocket and gRPC. A request and response cadence of roughly 10 transactions over two seconds is becoming 200 to 400 tokens per second at a 200-millisecond budget.
None of this shows up as a bandwidth problem; most enterprises have plenty of capacity. What CIOs see is a visibility and design problem. The network was built for a world of centralized, human-initiated, download-heavy traffic, and it is now carrying something structurally different.
Five Control Points the Network Must See
- Users - the employees and customers initiating requests are still the starting point for most policy today.
Agents - autonomous or semi-autonomous software that acts on a user’s or a business process’s behalf, often calling multiple models and tools in a single workflow without a human in the loop for each step.
Data - the enterprise information moving into prompts, retrieval pipelines and fine-tuning sets that needs to be tracked both inside a database and as it crosses a firewall.
Models - the specific inference endpoints being called, which could be a hyperscaler-hosted frontier model, a fine-tuned small language model, or a partner-hosted API, each with a cost, latency, and risk profile.
- Harness - the orchestration layer that wraps a model with tools, memory and permissions and turns it into an agent capable of taking action. The harness is where intent turns into behavior, which makes it one of the most important and least visible control points.
Traditional network and security tools were built to see users. The other four are largely invisible to SD-WAN, SASE, and conventional NetOps stacks today. The AI-grade network has to close this gap.
Proof Point: Two Moments of Truth for a National Bank
Consider this anonymized scenario seen at a large financial services team running a hub-and-spoke network. The bank has thousands of branches and ATMs connecting back to a small number of owned data centers through a mix of MPLS, dedicated internet and fixed-wireless backup over several carriers for redundancy.
Historically this architecture was sufficient because their traffic was predictable and applications lived where the network expected them. Two initiatives are stress-testing that configuration.
- Video banking at the ATM. Live video assistants at ATMs and interactive teller machines backed by a generative AI assistant (often working alongside a remote human teller) for account questions, disputes, and basic lending inquiries. The session is a sustained, symmetric, jitter-sensitive video stream running in parallel with a token-level exchange such as speech to text, a retrieval lookup against the customer’s account, a model-generated response, speech back out –all inside the same few seconds.
- Fast loan decisioning. Small-business and consumer loan decisions need to move from days to minutes. This requires an agentic workflow that fans out to several distinct model and API calls, often across more than one cloud inside a tight SLA before the customer sees an answer.
Three Layers, One Architecture
For the bank, they need a physical network that is strong and reliable for classic branch traffic, that can also identify if a degraded video call was a network problem or a model problem, or which of several parallel model calls is the one holding up a loan decision. Solving this requires three layers, each solving a distinct part of the problem, working together as one architecture rather than three separate products.
Diagram shows an AI-grade networking model composed of three layers: physical network, intelligent NaaS, and token controls.
Each layer working as a standalone fails on its own. The bank needs all three layers, operating as a single system and not three vendors that have to be reconciled, to that understand both the bytes (is the video session smooth, is the API call getting through) and the tokens (is the right model responding, and responding fast enough).
Why the Token Control Layer Must Stand Separate
Most large enterprises, including the bank in the example, don’t run a single-carrier network. Physical and NaaS layers are intentionally multi-vendor providers spread across branches and back-office locations for redundancy and cost reasons. That’s the right design choice at the connectivity layer, and it should stay that way so enterprises have freedom to choose circuit providers.
The token control layer is different. If model routing, token-level security, and cost governance are enforced one way on the AT&T-connected branches and another way (or not at all) on branches riding a different carrier, the bank ends up with exactly the gap it’s trying to eliminate, just distributed across more vendors.
A loan-decisioning workflow shouldn’t be governed, or troubleshot, any differently depending on which circuit happens to carry its traffic that day. Neither should the video assistant running at an ATM in one region versus another. That is why a token control layer has to be built to sit above the access layer and travel with the workload rather than with the pipe.
The token layer should see and govern AI traffic consistently whether it enters over owned fiber, a partner’s last mile, or a third-party carrier, while reserving the richest telemetry and tightest SLA for the access an enterprise chooses to own. One intelligence layer, many circuits underneath, not the other way around.
Where the Obvious Alternatives Fall Short
Given how much attention this space is getting, it’s worth being direct about two approaches enterprises are likely to encounter and why neither one, on its own, closes the gap.
Pure model routing. Software-only AI gateways can route calls by cost, latency, or accuracy but they have no visibility into the network underneath. They can tell you a call was slow; they can't tell you whether that's a model problem, a congested circuit or a carrier issue. Routing intelligence without network intelligence only optimizes half the problem.
Connectivity-only NaaS. Programmable, API-driven, on-demand bandwidth is now table stakes, not differentiation. It moves bytes efficiently, but SD-WAN and SSE were never built to reach models, agents, or tokens.. Traffic to and from hyperscalers and neoclouds still shows up fragmented and inconsistently secure, because the tools weren’t designed to know what that traffic actually is.
A third pattern worth mention: hyperscaler-native approaches that solve visibility and control inside a single cloud. These work well until an enterprise (like our bank) needs a loan-decisioning workflow to reach models running across more than one cloud. Then the single-cloud approach breaks, and the enterprise is back to stitching together per-cloud tools with no common view.
The Purview: What Enterprises Need Visibility and Control Over
Back to the bank’s network team. An AI-grade architecture gives them command over a specific set of concerns that didn’t exist in this form a few years ago:
First-mile vagaries. Branch and ATM connectivity is inherently variable: fixed-wireless failover, mixed carrier quality, differing installation timelines. That variability affects more uptime and determines how a live ATM video session stays smooth and whether the assistant behind it responds without interruption.
Last-mile fluctuations driven by intelligence, not just load. Where earlier routing decisions were based on congestion and static QoS classes, AI-era routing has to react to model and agent behavior in real time and adjust accordingly.
Control. Per-model and per-agent access control, policy set by workload, user, region, and data sensitivity, and the ability to contain an agent’s blast radius - governance the bank’s network team is asking for on both the video and lending use cases.
Visibility. AI traffic is inherently difficult to trace across models, agents, APIs, and retrieval flows. End-to-end telemetry from packets to tokens provides data-movement visibility and a single view that identifies the breakdown when a video call stutters or a loan decision is slow.
Security at network and the token level. Where traditional security focused on connections, AI-era security requires per-model and per-agent authentication and authorization, token-level inspection of prompts and responses, and data-loss prevention that understands what’s inside an inference call, not just that a connection was made.
Where This Leaves Us
The physical network and the programmable NaaS layer remain foundational but must advance. What’s now needed is the layer above them that understands workloads, agents, models, and tokens well enough to govern them, and that stays consistent no matter how many circuit providers sit underneath it.
As AT&T Business team heads to Gartner IT Symposium, I’m working on the final post in this series: what this architecture looks like as a product, the capabilities it delivers layer by layer, and the value it creates for enterprises putting AI to work in production.
Read more AT&T Business news