In version 1.0 of the Laravel AI SDK, classification becomes a capability of its own, alongside text, images, audio and embeddings. That is not a catalogue detail: it is the framework saying out loud something anyone running an AI product has already paid to learn — that "answer a closed question" and "generate something" are two different jobs, with different costs and different risk profiles. In Miraviso that split was not an optimisation that arrived later. It is the reason the product can claim that hair colour never leaves the device.
What it looks like in the SDK
The API is deliberately narrow. You ask a set of typed questions and get typed answers back:
use Laravel\Ai\Classification;
use Laravel\Ai\Classification\Boolean;
use Laravel\Ai\Classification\Choice;
$response = Classification::of($ticket->body)->questions([
'is_urgent' => new Boolean('Does this message convey urgency?'),
'department' => new Choice('Which team should handle this?', [
'billing' => 'Payments, invoicing, refunds',
'technical' => 'Bugs, outages, integrations',
'sales' => 'Pricing, plans, upgrades',
]),
])->classify();
There is a Score type returning 0.0 to 1.0, and a Str::decide macro for a single yes/no question with a certainty threshold. It runs on TypeSafe's Jev models and on OpenRouter; the announcement promises answers "in milliseconds at a fraction of the price" of a traditional LLM. I take that for what it is — a vendor claim, to be measured against your own workload — but the shape of the API interests me regardless of the benchmark.
The point is that the return type is closed. $response['department']->choice is one of the three values you declared, or it is nothing. No JSON to repair, no prompt politely begging for a one-word answer, no retry for when the model felt like explaining itself. Anyone who has shipped an LLM as a router knows how much defensive code disappears when the contract looks like this.
The three tiers I sort work into
In Miraviso — the virtual mirror for hair salons — every piece of work lands in one of three tiers, and the tier is decided before any code is written.
On the device. Everything that is geometry and segmentation. The colour try-on runs entirely on the tablet with MediaPipe on-device: the model finds the hair in the frame, the colour is composited locally, no frame leaves the hardware. This is not a cost decision, it is a product decision: it is a sentence I can say to a salon owner without an asterisk, and it is verifiable by watching the network.
Closed decisions. Questions with a finite set of answers — the ones that now have a dedicated capability in Laravel's SDK. In my case almost all of them sit on the back-office side and never touch an image: routing an incoming request, deciding whether a salon's note contains something to be treated as sensitive, deciding whether a piece of text should be shown to an operator. A generative model answers these perfectly well, and is the most expensive, slowest and hardest-to-test instrument I could possibly pick for them.
Actual generation. Exactly one thing: the haircut preview, which requires producing new, plausible pixels. It runs on Gemini via Vertex AI in an EU region, requires explicit consent, and is never written to disk. It is the only part of the pipeline that needs a large model — and also the only one that carries a privacy story I have to explain in full to the end user. As I wrote about tool approvals in the SDK, the boundary that matters sits before the model, not inside the agent.
Why the tier has to be decided up front
The temptation, once you have a working generative integration, is to use it for everything: the client is already there, the error handling is already there, one more question costs five minutes. That is precisely how a product accumulates generative calls nobody designed.
Three consequences I have watched happen in my own code, in order of how much they annoyed me:
- The data perimeter widens quietly. Every question handed to a remote model is one more piece of data leaving. If a classifier or a plain rule could have answered it, you widened the perimeter for nothing — and the perimeter is the thing you eventually have to describe in a privacy notice.
- Tests stop being tests. A closed question has a set of acceptable answers and you assert on it. A natural-language answer gets checked by another model, or by eye, and at that point the suite no longer tells you whether you broke something.
- Deprecations hurt more. I wrote about this already: when an endpoint is retired, the migration hurts in proportion to how many decisions you leaned on that model for. Closed decisions can be re-measured in half a day. Generation cannot.
The craft has not changed
Said without reverence for the historical moment: this is the same question we have been asking for twenty years in front of a slow query. Do I really need to call the service, or can I compute the answer here? The difference is that this time the call does not only cost latency — it costs money per request and it moves data across a boundary you promised to hold.
It is also why I keep hand-writing the logic games on this site with no libraries: when the minesweeper solver has to guarantee that a grid never forces a guess, that guarantee comes from a proof, not an estimate. Not every hard problem is a model problem — and telling which ones are not has become, over the last two years, one of the few skills that actually moves the cost of a product.
Which is the good news here: having classification as a separate capability makes that choice explicit in the code, instead of leaving it implicit in a prompt. A framework that forces you to declare what kind of work you are asking for is a framework that helps you not ask for the wrong one.