Telecommunications organizations are more and more seeking to AI to assist groups navigate extremely specialised domains, however generic fashions usually lack the industry-specific information wanted to know telecom networks, requirements, and operations. To handle that hole, AT&T created their Open Telco (OTel) fashions, the subsequent technology of telecom-focused AI designed to carry deeper telecommunications experience into AI techniques. Constructing OTel2.0 required greater than coaching a big language mannequin, it mirrored a broader challenge many organizations face: how you can construct domain-specific AI techniques at scale whereas balancing price, efficiency, and operational complexity. Price administration shortly grew to become a key consideration. To proceed advancing telecom-focused AI, AT&T wanted a platform able to supporting OTel2.0 growth at a wholly new scale.
The place groups beforehand needed to personal and handle deployments, infrastructure, and the related operational overhead, Foundry Managed Compute supplied a extra streamlined option to entry devoted graphics processing unit (GPU) capability. This transformation requires greater than highly effective fashions; it requires the flexibility to scale with out compromising price, flexibility, or efficiency.
Utilizing Microsoft Foundry Managed Compute, AT&T was in a position to experiment throughout a number of open fashions, optimize workloads throughout totally different GPU architectures, and course of large volumes of telecom knowledge all inside a unified platform. The end result was an AI growth surroundings able to supporting trillions of tokens whereas giving groups the pliability to iterate, optimize, and innovate sooner.
Mannequin alternative meets infrastructure flexibility
Constructing OTel2.0 required flexibility throughout each fashions and infrastructure. Quite than standardizing on a single mannequin, AT&T adopted a multi open-model technique. Open fashions have been central to AT&T’s strategy as a result of they supplied the pliability to work with authorized telecom knowledge, tailor the workflow for domain-specific mannequin growth, and assist large-scale experimentation with better management over price and deployment technique. By way of Microsoft Foundry, the crew deployed a number of fashions from the Hugging Face assortment, together with Phi-4, OSS-120B, and Gemma-4, to assist totally different levels of growth, from artificial knowledge technology and knowledge preparation to reasoning-intensive workloads and broader mannequin growth efforts. Phi-4 performed a major function on this course of, processing greater than 700 billion tokens a month as a part of the broader knowledge preparation and coaching workflow for OTel2.0.
Each firm on the earth must construct its personal AI, and that’s solely attainable with open fashions and open supply. AT&T is championing this imaginative and prescient, constructing on open fashions like Phi-4 and Gemma, and giving OTel again to the neighborhood as a telecom AI basis others can construct upon. Microsoft Foundry makes this sensible at scale, bringing the newest open fashions from the Hugging Face assortment along with AMD and NVIDIA GPUs in a single place, so groups can decide the appropriate mannequin and the appropriate {hardware}, then deploy in hours as a substitute of weeks.
—Jeff Boudier, Vice President of Product, Hugging Face
Growing OTel2.0 additionally required infrastructure able to working at telecom scale. AT&T used roughly 530 GPUs by means of Microsoft Foundry Managed Compute spanning a number of GPU architectures together with 430 AMD Intuition™ MI300X GPUs. This heterogenous strategy gave AT&T extra flexibility in how fashions have been deployed and optimized as necessities developed.
| Mannequin | Instance workload |
|---|---|
| Phi-4 | Round 700B tokens a month for knowledge preparation and artificial knowledge technology |
| OSS 120B | Larger-reasoning workloads |
| Gemma 4 | OTel2.0 growth workflows |
This flexibility illustrates a broader pattern throughout AI growth. Organizations more and more want platforms that enable them to decide on the appropriate mannequin for the job, optimize for price and efficiency, and scale workloads with out rebuilding operational environments. Microsoft Foundry brings mannequin alternative, infrastructure flexibility, governance, and operational scale collectively in a unified platform that helps these necessities.
Past flexibility and value, deployment pace is a crucial issue for a lot of AI initiatives. As workloads broaden and new fashions are evaluated, the flexibility to entry GPU capability shortly permits groups to maneuver from experimentation to execution sooner with out prolonged provisioning cycles. With Foundry Managed Compute, AT&T may deploy and scale fashions in days reasonably than ready weeks for infrastructure to turn into accessible, serving to speed up growth timelines and keep momentum throughout OTel2.0 growth.
Optimizing price with out limiting innovation
As AI workloads develop, economics turn into as essential as mannequin efficiency. For AT&T, one of many major goals was to decrease AI mannequin consumption prices whereas persevering with to drive significant enterprise worth by means of AI-powered innovation. Through the use of open fashions on Microsoft Foundry Managed Compute, AT&T was in a position to assist large-scale knowledge preparation and mannequin growth utilizing a unique financial mannequin constructed round devoted GPU infrastructure and open-model flexibility.
The influence grew to become clear at scale. In assist of OTel2.0, AT&T processed roughly 1T tokens, consisting of uncooked paperwork from GSMA supplemented by artificial knowledge generated. Producing the info utilizing open-source fashions like Phi-4, served by Microsoft’s Foundry Managed Compute, saved tens of thousands and thousands of {dollars} versus utilizing frontier fashions. This allowed groups to spend money on larger-scale experimentation and growth whereas sustaining a concentrate on enterprise worth and operational effectivity.
| Metric | Worth |
|---|---|
| OTel 1.0 Downloads | Over 25M |
| GPUs Used By way of Foundry Managed Compute | About 530 |
| Tokens Processed for OTel2.0 | About 1T |
| Tokens Skilled for OTel2.0 | About 400 B |
| Fashions used to coach OTel | Phi-4, OSS 120B, Gemma 4 |
If you end up processing lots of of billions of tokens, infrastructure turns into a part of the issue you remedy. Foundry Managed Compute gave us entry to GPU capability at scale so our groups may concentrate on advancing OTel2.0 as a substitute of managing infrastructure.
—Mark Austin, Vice President, Information Science and AI at AT&T
At this scale, infrastructure is now not merely a deployment consideration. It turns into a strategic element of AI growth.
Accelerating the subsequent wave of production-scale AI
OTel 2.0 demonstrates how organizations can mix open fashions, scalable infrastructure, and area experience to construct production-ready AI techniques. By matching totally different fashions to totally different workloads and optimizing infrastructure for price and efficiency, AT&T was in a position to course of trillions of tokens whereas sustaining operational effectivity.
As organizations transfer from AI experimentation to manufacturing deployment, they more and more want the pliability to decide on the appropriate fashions, optimize infrastructure, and scale effectively. Microsoft Foundry and Foundry Managed Compute assist assist that transition by bringing these capabilities collectively in a unified platform.
Be taught extra
Discover session matters from AMD’s Advancing AI:
