In-Depth

Why Tokenomics Is More Than Just Counting Tokens

A few weeks ago, I had the pleasure of attending Splunk's annual conference in Denver. One of the big topics of conversation at the event was something Splunk refers to as tokenomics. As the name suggests, tokenomics is essentially the idea of managing the costs associated with AI tokens.

On the surface, that concept is deceptively simple. The logical assumption is that tokenomics is all about controlling costs by counting tokens. In fact, one could argue that, to date, this has been the primary way organizations have attempted to control their AI costs. A company might say, "We consumed 10 million tokens last week. We know the cost per token, so we spent X dollars on AI."

In reality, though, the cost-per-token metric is almost meaningless because it assumes that all tokens are equally valuable.

Splunk's approach to tokenomics is to use its footprint within the enterprise to create actionable spending insights. The basic idea is that if you can attribute costs to specific users, teams, models, tools and workflows, then it becomes easier to see where token-related costs are coming from and to create a spending forecast before the bill arrives.

All of that is undeniably useful, but I think it only begins to scratch the surface of the broader discussion that needs to happen around AI spending. Cost forecasting is valuable, and visibility into token consumption is critical, but the thing that might ultimately matter more is cost optimization.

Cost optimization does not always mean cutting costs. Instead, it can mean focusing your spending in a way that increases business value. Let me give you an example.

Obviously, no AI is perfect, but some models are inevitably going to do a better job than others for a specific use case. So let's pretend that there are two models that are both capable of performing a specific task. The first model costs 10 cents per operation but is successful 50% of the time. The second model costs 20 cents per operation but is successful 99% of the time.

In terms of raw dollars spent, the second model is far more expensive. It costs twice as much to operate. But if you examine the cost per successful outcome rather than just looking at the raw cost, the difference becomes much smaller.

The first model has a cost of 20 cents per successful operation. The second model's cost is about 20.2 cents.

Admittedly, that 2-cent difference can become significant if the model is operating at a large enough scale. However, it is also worth considering operational efficiency.

To see why this matters, let's assume that both models deliver a comparable level of performance. The first model, though less expensive from a token standpoint, is going to incur additional costs in the form of retrying unsuccessful operations. These retries are going to consume network bandwidth, storage IOPS, compute resources and other infrastructure.

Assuming that the model is being hosted in the cloud, this overhead could easily eliminate the apparent cost advantage of operating a less expensive model. In other words, token cost matters, but there are other factors that must also be considered.

And, of course, the unsuccessful outcomes from the previous example may carry additional costs. Let's assume that the AI doesn't do anything catastrophic, such as making a bad database modification or an erroneous financial decision, but it does require a human to actively take some sort of corrective action. That human in the loop costs money.

To put it another way, the workload cost is not determined solely by token cost. The actual cost is also based on risk.

This is where risk-based tokenomics comes into play. The economics of an AI workload can be thought of as consisting of the inference cost plus the probability of a failure multiplied by the failure cost.

Of course, we also have to talk about the economics of autonomy because those economics can be surprisingly counterintuitive.

Over the last few years, we have seen organizations progress from treating AI as an assistant to adopting supervised AI agents and ultimately to experimenting with autonomous agents. The assumption has always been that autonomy reduces labor costs, scales well and is absolutely the path forward.

Under that assumption, an agent that costs $1,000 per month but operates autonomously would appear to be more economically viable than an agent that costs $100 per month but requires constant human supervision.

Autonomy can indeed decrease human labor costs, but it can also increase both inference costs and monitoring costs. Perhaps more importantly, autonomy can potentially increase the blast radius when things go wrong, since the agent in question is making consequential decisions without human involvement.

That raises an arguably more important economic question: At what point does additional autonomy stop being economically rational? At what point does the cost of an agent making a bad decision outweigh the cost of requiring a human to approve certain decisions?

As models become more capable, there is little doubt that we will see organizations gravitating increasingly toward AI agent autonomy. Even so, I do think there is a viable way to strike a balance between token costs, performance and risks.

Rather than just counting token consumption, which is what many organizations are doing today, it may make more sense for organizations to treat the model selection process like an economic control plane.

Think back to the example that I gave earlier, in which two models were each capable of performing a task. One was more accurate, while the other was less expensive to use.

In the real world, there may be more than two models that can handle a particular job, and each of those models will probably have its own tradeoffs regarding performance, quality, latency and cost. As a result, we may begin seeing organizations handling AI model selection in a way that somewhat resembles a sophisticated cloud scheduler. A particular task might have a set of rules governing which model should be used in a given moment. That list of rules may look something like this:

  • Under normal conditions, use Model A.
  • If your confidence in the output drops below 90%, escalate to Model B.
  • If latency becomes critical, switch to Model C until the demand subsides.

The important takeaway from this is that the system would not necessarily be choosing the cheapest model. It would be continuously trying to strike the right balance between cost, performance, risk and business requirements.

All of this is to say that while tokenomics is destined to become a major trend, the goal of tokenomics should not be to reduce spending by using fewer tokens. Instead, the goal should be to use the right number of tokens to create the desired business outcome without unnecessarily wasting tokens in the process.

About the Author

Brien Posey is a 22-time Microsoft MVP with decades of IT experience. As a freelance writer, Posey has written thousands of articles and contributed to several dozen books on a wide variety of IT topics. Prior to going freelance, Posey was a CIO for a national chain of hospitals and health care facilities. He has also served as a network administrator for some of the country's largest insurance companies and for the Department of Defense at Fort Knox. In addition to his continued work in IT, Posey has spent the last several years actively training as a commercial scientist-astronaut candidate in preparation to fly on a mission to study polar mesospheric clouds from space. You can follow his spaceflight training on his Web site.

Featured

Subscribe on YouTube