Stacklet targets AI token spending with policy controls
Stacklet's early-preview Token Custodian brings policy controls to AI spending, raising a question for CIOs: Who decides which work gets access to more expensive AI models?
The ability to track and manage costs is a basic IT function, but AI has made it increasingly difficult. The complexity of token consumption and how it relates to governance means that simply setting a budget isn't enough.
Stacklet's Oct. 6 launch of Token Custodian shows where that problem is heading. The early-preview product is designed to go beyond tracking AI spending to enforce usage policies, route work among models and handle exceptions. Stacklet describes it as a control plane for AI token usage.
Enterprises faced a similar challenge when cloud costs began to spiral amid weak controls and governance, helping establish FinOps as a practice for managing cloud spending and maximizing its value. In the AI era, similar ideas are moving into tokenomics, or the management of AI token use and cost.
Cloud Custodian, an open source governance-as-code project created at Capital One, became an early tool for enforcing cost, security and compliance policies across cloud infrastructure. Members of its core team later founded Stacklet, which built its business on cloud governance.
Stacklet says customers asked for the controls they use in the cloud to be applied to AI. The comparison has limits, though.
Traditional cloud cost allocation often starts with resources, accounts or services that have identifiable owners. Model API charges can accumulate through requests made by employees, applications and agents, making attribution more complicated.
The FinOps Foundation's FOCUS 1.5 working draft, scheduled for ratification Dec. 3, includes a Principal ID field intended to identify the person, service account or other entity behind a charge. Stacklet is tackling a related problem with Token Custodian, aiming to connect AI spending to the teams, applications and agents generating it, then use that information to enforce spending policies.
"When you think about cloud, that's supporting the products or services that the organization is bringing to bear for its own customers or employees," Travis Stanfield, CEO of Stacklet, told TechTarget. "With AI, you're engaging like all of the knowledge workers in the organization to essentially be more productive."
What Token Custodian does
The major AI vendors generally provide some degree of budget and cost control capabilities. Stanfield said many of those controls are rudimentary, an on-or-off choice, and that enterprises using several providers would have to reapply them in each provider's tools. According to Stacklet, Token Custodian is intended to address both gaps with one policy layer across providers and graduated responses in place of a hard stop.
Attribution. Stacklet says Token Custodian is designed to trace spend to the team, project, application and cost center behind it, and to the agent involved.
Policy. Stacklet says budgets and usage policies can be set by team, project or environment.
Responses. When a team nears a limit, the proposed controls provide alternatives: "Do we need to manage an exception? Do we need to perhaps route the model usage to something that's more economical?" Stanfield said. Options include a notification, an exception request, a shift to a more economical model or a stop.
Exceptions. Stacklet says exceptions can be time-boxed, for example to a week or two, and then sunset.
One key challenge that organizations are coming to terms with is trying to determine the ROI from token spend. Stanfield said Token Custodian is not an ROI calculator at this stage but could serve as a building block for those calculations.
Enforcement is still early
The Tokenomics Foundation's Sept. 23 Tokenomics report, based on 472 responses across 11 industries, found that organizations primarily rely on spending caps and token limits to govern developer productivity costs. But its broader findings point to a bigger challenge: Proving AI's business value (43%) and tracking spending across systems (27%) were among respondents' leading concerns. Stacklet is a founding member of the foundation.
Stanfield said enforcement is necessary in the token world because costs are rising so quickly. Visibility and observability alone cannot keep pace as agents automate more work, he said.
In practice, however, enforcement remains a step behind.
"The broad challenge right now is that most AI cost control is still manual and after the fact," Lindbergh Matillano, director of cloud and AI optimization at Avalara, told TechTarget. Avalara is a Stacklet customer and has plans to try out Token Custodian in the future. He said teams have gotten good at seeing their AI spend, while governing it is the harder, less-solved part. Real-time control that does not hard-block engineers is difficult with current tools, he said.
Where controls exist, they may still be blunt instruments.
You do not want FinOps or platform engineering to lower software quality by restricting model access.
Torsten Volkprincipal analyst for application modernization, Omdia
"Caps, approvals and model routing all are important, but they do not allow to optimally align token use with business importance," Torsten Volk, principal analyst for application modernization at Omdia, a division of Informa TechTarget, said.
Volk said Token Custodian is designed to detect which application, developer or project team is responsible for a certain call to a specific LLM and enforce policies based on that. For example, a policy could allow a code function serving key customers to use a frontier model.
Another function from that same developer might enable features of a non-critical internal app and therefore have its LLM calls routed to more cost-efficient models. Routing could also reflect the capabilities a task requires: Simple categorization work could go to a potentially lower-cost specialized decision model such as Jev, while demanding code creation tasks go to higher-end LLMs.
"Token Custodian enforces the policies in real time that add this additional level of business alignment to model routing," he said.
Who sets the policy
A key question for organizations is who sets the policy for AI token consumption.
Stanfield said every FinOps practitioner he talks to is being asked to get involved in their organization's AI challenges, though they may not be the final arbiter of the budget. He said he tends to see three groups involved. They are FinOps, AI and machine learning platform leads and the teams closest to the developer experience.
Who should control AI model spending?
AI cost controls do more than limit spending. They can determine which models employees and applications use, when access is restricted and whether an exception is justified. Those decisions can affect software quality and business outcomes, so setting the rules requires input from several teams.
Volk said the call should not rest with FinOps or platform engineering alone.
"You do not want FinOps or platform engineering to lower software quality by restricting model access," he said.
In his view, those groups would not even remotely have the knowledge to decide which models are appropriate for a certain development workflow or software feature.
"Reining in token cost really is a multi stage process, where first you need to understand what everything costs," Volk emphasized. "Then, the FinOps guys can go to the development lead and ask if there are options for cutting frontier model use."
Volk noted that there might be a lot of low-hanging fruit and the development lead might be happy to assign a lower tier version of a frontier model to a certain set of development tasks.
"This initial conversation between FinOps and developers alone starts a culture where project teams are more aware of Token cost," Volk said. "They will often be able to cut down significantly on Token usage without any significant deterioration of code or agents that are part of their product."
Cheaper tokens aren't the goal. Better value from them is.
Lindbergh Matillanodirector of cloud and AI optimization, Avalara
"Even on the token side, attributing spend to the right team and workflow is harder than it should be, and a lot of teams are still doing it manually," Matillano said. He said costs beyond tokens, such as infrastructure and human review, also exist, but that control has to start at the token and usage layer.
Cost controls can pull teams toward the number they can see and that might miss the point.
"The trap across the industry is cutting the cost you can measure and losing the outcome that mattered," Matillano said. "Cheaper tokens aren't the goal. Better value from them is."
Sean Michael Kerner is an IT consultant, technology enthusiast and tinkerer. He has pulled Token Ring, configured NetWare and been known to compile his own Linux kernel. He consults with industry and media organizations on technology issues.