July 23, 2026

The Tokenomics Equation: Balancing Cost and Performance

How we’re getting more value from AI by blending proprietary and open-source technologies

Andy Markus
Andy Markus Chief Data and AI Officer, AT&T

Artificial intelligence is a cornerstone of telecommunications’ future. Because we work with some of the most complex data sets in the world, we have a real opportunity to transform how we automate complex network operations and deliver more personalized customer experiences. But we can’t lose sight of one of our most critical responsibilities: using AI efficiently.

We’re focused on efficiency in two primary ways: we’re using a proprietary cache-aware router to select the most cost-effective models for each task, and we’re training more open-source models to address telco-specific needs for accuracy. And both of these approaches are being actively used in production to drive real results, not just theoretical ones.

Cracking the Cost Challenge: A New, Intelligent AI Gateway

Running advanced AI models can be expensive, especially at AT&T’s scale: an average of 45 billion tokens per day.

People often default to the latest and greatest models, but only a small percentage of the tasks we run require that level of sophistication. Many can be handled by lower-cost models without sacrificing performance.

This is where our data science team shines. We’ve built an AI Gateway that goes beyond simple model selection: it uses cache-aware routing to intelligently match each task to the most cost-effective model—without compromising on quality. It can even change models partway during multi-turn sessions.

At each turn, the gateway weighs speed and cost with the expected quality of the output, then routes the prompt to the best model.

The result: we’re reducing AI costs as much as 90%, already saving millions.

Answering the Accuracy Challenge: An Expanded Library of Telco Models

Efficiency only matters if accuracy keeps pace. The scale and complexity of our networks demand solutions that are both powerful and safe. We’ve been vocal about our work to open-source models that understand telco’s unique challenges.

Now we’re entering into the next phase of open telco AI, launching the newly trained OTel 2.0 model, and proving the flexibility of the hardware that runs the model.

AT&T is one of the first companies to use Managed Compute at scale through Microsoft Foundry, processing over 1 trillion tokens for training. With Managed Compute, we’re training dedicated models on powerful AMD hardware and data provided by GSMA.

We’ve post-trained OTel 2.0 using over 400 billion tokens on AMD GPUs.

We’re training these models to provide the telco industry with more flexibility and better optimization. They drive down token costs by providing model options that are suited to our tasks, and can run on right-sized hardware, either on premises or in the cloud.

It’s good for our business. But more importantly, it’s good for our customers.

Truly Sovereign AI = AI Gateway + Specialized Models + Governance + Security

Every part of the AI ecosystem needs to work together. The costs and the accuracy need to be combined with strong security, and robust governance. When all of those pieces work together, this sets up a company for a viable AI management system that’s been gaining stead: Models as a Service.

ModaaS provides end-to-end control and management of your AI stack — from governing the flow of proprietary data to trusted sources, to optimizing the costs of running AI at scale. It's a major lever for protecting enterprise IP, driving return on your AI investments, and transforming operations from the ground up. And it provides a pathway to truly sovereign AI.

I believe AT&T is building something truly unique by combining our telco expertise with our AI proficiency—and setting a foundation that can extend well beyond telecom.

Sometimes it takes outside-the-box thinking to find the right solution. That’s the kind of thinking AT&T puts into practice.