Skip to main content

        Enterprise AI should use the smallest, fastest and most economical model that can reliably deliver each business outcome—not chase frontier intelligence for every workload.

The Minimum Viable Model: Why Enterprise AI Should Stop Chasing the Frontier

Enterprise AI should use the smallest, fastest and most economical model that can reliably deliver each business outcome—not chase frontier intelligence for every workload.

Lately, much of the AI conversation has revolved around one question:

What is the best model?

GPT, Claude, Gemini, Chinese models… Every few months there is another release, another benchmark, another leaderboard and another model claiming the frontier.

Enterprises increasingly need to ask a different question:

What is the smallest, fastest and most economical model that can reliably solve this particular problem?

The other day I heard the term Minimum Viable Model, or MVM.

The principle is simple:

The best model for an enterprise workload isn’t necessarily the most intelligent model available. It’s the least expensive model that meets the required level of quality, reliability, latency, security and governance.

And as models improve, that threshold is moving down remarkably quickly.

Frontier performance is becoming a moving target

Last night I was playing with Qwen 3.8 27B on my 4090 desktop PC. Some say it’s better than Opus from six months ago. Is it really? Does it need to be?

Think about what that represents.

Capabilities approaching what we would have considered frontier class not very long ago can now run locally on a gaming PC.

If a smaller model produces the required result consistently for your specific business process, the additional intelligence of a frontier model may provide little or no additional economic value for that process.

Yesterday’s frontier capability is rapidly becoming today’s commodity capability.

Tokenomics changes the architecture

This is where enterprises need to start thinking seriously about tokenomics.

When experimenting with AI, token costs can look insignificant.

At enterprise scale, token economics becomes architecture.

Imagine an agentic workflow performing classification, extraction, summarisation, retrieval, validation and generation.

A single user action might trigger dozens of model calls.

Now multiply that by 10,000 employees.

Then multiply it again by hundreds of interactions.

Then introduce autonomous agents that operate continuously without a human initiating every request.

Suddenly the question isn’t:

Which model gives us the highest benchmark score?

It becomes:

Why are we paying frontier model prices for a task that a much smaller model can perform correctly?

Consider document classification.

You probably don’t need the most capable reasoning model available.

Extracting structured information from an invoice?

Probably not.

Summarising a support ticket?

Probably not.

Routing a request to the correct internal system?

Again, probably not.

Solving a complex architecture problem with ambiguous requirements, reasoning across multiple systems and making consequential recommendations?

Now the frontier model may be exactly what you want.

MVM isn’t about choosing one cheaper model for everything.

It’s about choosing the right level of intelligence for each workload.

The enterprise AI stack should look like a pyramid

Minimum Viable Model pyramid showing efficient, strong and frontier model tiers, plus a dynamic model-selection and escalation path

At the bottom of the pyramid are high-volume, predictable tasks.

Classification. Extraction. Transformation. Basic summarisation. Simple RAG. Entity recognition. Routing.

These should use small, efficient models wherever their quality is sufficient.

Move upward and you encounter tasks requiring greater reasoning, longer context, multimodality or sophisticated tool use.

Use a stronger model.

At the very top are the genuinely difficult problems: complex reasoning, ambiguous decision-making, sophisticated coding, planning and autonomous multi-step work.

That’s where frontier models earn their premium.

The architectural mistake is putting the entire pyramid on the model at the top.

Model selection should become dynamic

The next logical step is to stop thinking about model selection as a static architecture decision.

Instead of:

Application → Model

I expect more enterprise architectures to look like:

Application → AI Gateway → Model Router → Minimum Viable Model

The router evaluates the task.

Simple request? Small model.

Large document summarisation? Efficient long-context model.

Sensitive internal workload? Private model.

Complex reasoning? Frontier model.

Low-confidence result? Escalate to a stronger model.

Critical decision? Potentially use multiple models or introduce human validation.

This creates an intelligence escalation path.

Start with the Minimum Viable Model.

Escalate only when necessary.

The metric that matters isn’t intelligence. It’s useful intelligence per dollar.

The frontier will continue moving.

Models will continue getting larger, smarter and more capable. That research matters enormously.

But enterprise architecture has a different objective.

Enterprises don’t win because they consumed the smartest tokens.

They win because they transformed those tokens into business outcomes efficiently.

So perhaps the most useful question for enterprise AI isn’t:

What is the best AI model available?

It’s:

What is the Minimum Viable Model for this workload?

Once organisations start asking that question across thousands or millions of AI interactions, model selection stops being an AI trend discussion.

It becomes architecture. It becomes FinOps.

Because at enterprise scale, intelligence isn’t free.

And paying for intelligence you don’t need isn’t innovation.

It’s waste.