Commentary by Ethan James Farrell, Founder, Sensus InVista

Following on from the last post, where I detailed increasing benchmark performance against the cost of API access to AI models, what that demonstrated — at least across the examples I gave — was that the increase in benchmark performance versus the increase in API cost shows models having only incremental improvements, whilst the cost of providing these models is clearly significant for the provider.

In some cases, the API cost per million tokens has doubled versus a very minor increase in benchmark performance.

That’s obviously not the case across the board. Google Gemini showed an increase in benchmark performance with quite significant decreases in API service costs. However, this is more likely to be a strategic move due to the ever-changing landscape and competitive pressure.

What is particularly interesting in Google’s case is the scale of investment sitting behind that pricing. Alphabet’s second quarter of 2026 was its first negative free-cash-flow quarter since going public. It generated approximately $39.1 billion in operating cash flow whilst spending $44.9 billion on capital expenditure, resulting in approximately negative $5.8 billion in free cash flow. Its expected capital expenditure for 2026 has now risen as high as $205 billion, compared with $91 billion spent in 2025, with much of that expansion associated with the infrastructure required for AI.

Sensus InVista infographic showing Alphabet Q2 2026 operating cash flow of $39.1B, $44.9B in capital expenditure, projected 2026 CapEx of up to $205B, and negative $5.8B free cash flow.
The Mounting Cost of Unchecked Scale — Falling API prices do not necessarily mean falling AI delivery costs. Alphabet’s Q2 2026 figures illustrate the extraordinary infrastructure investment now sitting behind frontier AI.

So whilst Google’s API pricing may be moving downwards, that shouldn’t necessarily be taken as evidence that the underlying cost of delivering increasingly capable AI is doing the same.

Additionally, I detailed how significant performance increases were delivered through the use of agentic skills and configurations over fine-tuning. For me, this demonstrates that there is a case for innovation. Whilst that naturally requires some level of input in order to attain or achieve the performance increase on a task-by-task, case-by-case or domain basis, I do believe there is a case for providers to consider alternative models of operation or model availability.

A case for alternative model availability

There’s potential, I feel, for there to be licensing of models for local devices from big providers.

At the moment, whilst OpenAI has GPT-OSS and Gemini has Gemma, the core and best features of the major providers are generally only available through models accessed through their platforms.

Yes, naturally, they tend to be larger models, so on-device operation is somewhat out of scope for those models in particular. However, to reduce computing costs and computational load requirements, would it not make sense for something like GPT-OSS to be the in-app model, with calls being sent to higher models as required based on usage?

Would it not make sense for time to be spent improving the abilities, capabilities and technologies which go into quantisation, or the way models are trained, so that on-device becomes a preference — or the preferable modus operandi?

Three-stage Sensus InVista diagram showing an on-device AI model, a quantisation and security layer, and selective cloud handoff for computationally demanding tasks.
Re-architecting Model Availability — Local first, cloud when necessary: use existing consumer hardware for routine AI workloads and reserve hyperscale compute for the tasks that actually require it.

Many modern devices, particularly current smartphones and most consumer computers purchased within roughly the past three years, are already able to run some form of on-device model.

The significant investment being applied to the expansion of data centres — which are increasingly affecting communities and the perception of AI companies — could therefore be offset and reapplied towards more innovative approaches, harnessing the power of devices which are already in the hands of millions.

Providers could, for example, enact a licensing model. The model would be downloaded within an application on the device. I’m sure sufficient time and effort could be spent ensuring there is a security layer protecting the proprietary nature of the training techniques and perhaps the corpus of materials used to train the model — albeit that being a contentious issue in itself.

If those models could be secured in a way that allowed them to be distributed through application interfaces such as the App Store or Google Play, whilst ensuring the integrity of the model’s delivery to the device and its operation on-device through approved manufacturers and hardware, would it not make sense for that to be explored?

Compute requirements could potentially be offset massively, alongside addressing data-protection risks which are ever increasing and imposing increasingly negative consequences.

Sensus InVista infographic exploring how local-first hybrid AI deployment could improve environmental impact, public perception and user satisfaction.
Reinterpreting Deployment — Moving computation closer to the user is not simply a technical choice. It could also influence environmental impact, public trust and the quality of the everyday AI experience.

The wider case against continued expansion

I also believe that the race to AGI is flawed, in that I don’t believe human ingenuity could ever be replaced by technology.

The continued expansion of AI infrastructure — and the plans for expansion in both the short and long term — are increasingly creating negative connotations around AI. That comes alongside recent security risks which have been identified within organisations and, particularly, by manufacturers themselves, as well as calls from more than 1,100 professionals for tighter regulation.

That’s not to say that changes in the way these models are made available within applications or to users would not impose significant costs.

However, with the recent mass layoffs across the globe of highly specialised, highly capable software engineers and software coding professionals, there is a workforce which could be tapped into should there be a desire to change the direction of AI towards an approach that gives greater consideration to environmental, user and worker impact.

In terms of economics, naturally there would be a cost to applying a different approach to this area and applying workforce and resources towards the development of those approaches.

However, I think the investment cost could be much less than the cost of continued infrastructure expansion when considered not only in monetary terms, but also through environmental factors, public perception and the opportunity for AI leaders to demonstrate a clear commitment to addressing ongoing concerns.

Naturally, there is a need for — and benefit from — the high levels of performance that larger models can provide.

However, your average individual using AI for relatively simple tasks, or even an employee using it to help with their workflow, is unlikely to need the raw processing power and intelligence required for coding, complex mathematical operations or perhaps research involving very large datasets.

For those tasks, a local, powerful, proprietary on-device model which hands off, when required, to other models for things such as media generation and very high-demand tasks would be an appropriate area for attention.

Sensus InVista comparison matrix contrasting continued AI data-centre expansion with on-device innovation across capital demand, compute source, model strategy and data risk.
Expansion vs. Innovation — The alternative to continued infrastructure growth is not less capable AI. It is a different allocation of intelligence: compressed locally, distributed widely and escalated to frontier compute only when required.

Public perception and meaningful adoption

The public perception of AI is increasingly leaning towards the negative, in part because of the fear of replacing jobs.

But the increasing publicity and attention around data centres brings additional concerns, particularly around the harm they potentially pose to the health and well-being of communities where they are placed.

That disruption, tied with a fear of job replacement, will only foster a perception of fear or disdain if people are not also able to understand and experience the clear benefits that AI could provide.

That, in itself, risks becoming a limiting factor towards adoption — or, more importantly, meaningful adoption.

The question, then, is whether the next stage of AI development really needs to be defined primarily by expansion: ever-larger models, ever-greater compute requirements and ever-more infrastructure.

Or whether some of that effort could instead be directed towards innovation in how the technology we already have is trained, compressed, distributed and used — harnessing computing infrastructure which already exists in the hands of millions of people, whilst retaining access to frontier-level capability when it is actually required.

Further Reading:

Leave a Reply