Parallel Works
Back to Newsroom
podcastsactivateaihybrid multi cloud

Full Tech Ahead with Amanda - Control Your AI Spending

Share:

In this episode of "Full Tech Ahead," host Amanda Razani interviews Matthew Shaxted, CEO of Parallel Works. The conversation centers on a major obstacle facing enterprises today: skyrocketing token consumption and the ballooning costs of using frontier AI models.

Shaxted explains that opening up unrestricted API access to hundreds or thousands of users leads to rapid budget depletion, citing recent industry examples like Uber. Drawing a parallel to the high-performance computing (HPC) and cloud migration trends over the past decade, Shaxted predicts a cyclical shift: while firms currently rely heavily on public cloud endpoints, economic pressures and massive utilization rates will drive them to bring data workloads back on-premise using increasingly powerful open-weight models (like the recently released GLM 5.2).

To combat initial adoption chaos, Parallel Works offers a computing control plane called Activate, providing a single pane of glass to enforce visibility, tracking, and strict "token budgets" that automatically deny requests once expenditure thresholds are met.

Key Quotes

"Unless [token usage] is thought about in the very beginning in terms of how are you going to control and monitor token usage... it really becomes a big problem. We've seen the Uber story recently, where you burn through the entire budget in a few months."

"Having strong visibility in where the tokens are going... make sure from the very beginning you have visibility into who's doing what, because that's going to start growing very quickly."

"As soon as it basically is out of budget, it will deny the request until you get more allotted. That's exactly what's happening."

"As the open weights get better and better—which I think we're starting to see with GLM 5.2 coming out recently—you can start running open weight models for certain classes of things with much more predictable cost."

Takeaways

Establish Financial Guardrails Early: Unrestricted enterprise AI access creates a cash burn. Implement central governance and a "computing control plane" from day one to enforce hard spending limits (token budgets) at the user or team level, preventing unexpected tech invoicing.

Prepare for the On-Premise AI Hybrid Shift: Much like cloud computing evolved, enterprise AI will hit a baseline utilization rate (e.g., 8k/80% load) where renting per-token public APIs becomes economically unsustainable. Companies should plan a hybrid stack that shifts routine tasks to dedicated on-premise infrastructure running open-weight models to cut costs up to 6X.

Incentivize Token Awareness: End-users rarely optimize resource usage unless faced with explicit constraints. Providing visible token limits encourages developers and practitioners to delegate simpler, mundane prompts to lighter, less expensive internal models rather than burning resources on premium frontier labs.

Architect for Agentic Workloads: With the industry shifting toward massive agentic systems where hundreds of autonomous agents execute tasks 24/7, compute demands are projected to scale up to 1,000X. Managing this scale without breaking corporate cost structures requires unified virtualization and gateways across cloud and hardware assets.