In short
- The problem. When you add "a bit of AI" to business software, every question ends up going through a remote generative model. It costs on every request, widens the perimeter of data leaving the company and ties the product to a vendor's calendar.
- What I did. Before writing code I split the work into three tiers: what runs on the device, closed-answer decisions, real generation. Sensitive data went where the server cannot read it.
- The result. In Miraviso, my product, the part where AI works on customers' images never leaves the tablet. The part using a remote model is a single one, declared, and runs only with consent.
The context
Miraviso is a B2B SaaS for hair salons: my own product, in production, with a pilot programme open to the first salons. On the tablet it changes hair colour in real time and generates a photorealistic haircut preview in a few seconds. Underneath: FastAPI, PostgreSQL, Flutter, MediaPipe, Gemini on Vertex AI in the EU region, Docker, Caddy and Stripe.
Why a case study about a salon, if your company makes components, runs a hotel or ships goods? Because the questions are the same, and here I can show the details, which I would never do with a client's software. Where is an AI model really needed? What does it cost? What data leaves the company? What happens when the vendor retires the model?
The problem
A colour consultation runs on words: "a warmer brown", "blonde, but not too much". The client can't picture the result and falls back on the safe service. AI could help, but under three constraints that apply to any small business.
- The data. Notes about clients may contain allergies and scalp conditions: data bordering on health data, about people who never opened an account. In your company it may be orders, price lists, shifts, technical drawings.
- The cost. A generative model answers almost anything well, but for closed-answer questions it is the slowest and most expensive tool.
- The dependency. A preview model can be switched off by the vendor with about thirty days' notice. It just happened to another Gemini model, and I wrote about it in Thirty Days to Migrate.
What I did
Three decisions, all taken before writing the code.
1. Three tiers of work
Every task lands in exactly one tier, and the tier decides where it runs and what it costs.
| Tier | What goes there | Where it runs | Why |
|---|---|---|---|
| On the device | Finding the hair in the image and applying the colour | On the tablet, with MediaPipe | No frame leaves the device, and you can verify it by watching the network traffic |
| Closed decisions | Routing an incoming request, telling whether a note is sensitive, deciding whether a text should be shown to an operator | A classifier or a rule, not a generative model | The answer comes from a finite set: cheaper, faster, testable with assertions |
| Generation | The haircut preview, which needs new, plausible pixels | Gemini on Vertex AI, EU region, with consent, never written to disk | It is the only piece that really needs a big model |
The full reasoning, with code, is in the article When You Don't Need an LLM.
2. Sealed envelopes for sensitive data
For sensitive notes only, the encryption key is born on the salon's device and never reaches the server, which stores envelopes it cannot open. It doesn't apply to the whole product: calendar, appointments and invoicing are still processed server-side, as in any ERP. I priced in the costs: no server-side search on that field, no shortcuts for support, and key recovery to be designed with care. Details in the article The Server Can't Read It.
3. The model behind a boundary
A single place in the code knows which model is called and how. For every capability bought from a vendor there is an answer to the question "how do I replace it?". Preview images are not kept, so tests of a new model run on a set of images collected on purpose, with consent, outside production traffic.
The result
- The colour try-on never leaves the tablet.
- The haircut preview is the only step using a remote model: EU server, with consent, never written to disk.
- Sensitive notes sit on the server in a form the server itself cannot read.
- The colour try-on depends on no vendor: the model lives in the app and nobody can switch it off remotely.
- The product is in production on its own infrastructure, with an open pilot programme.
What I learned
- The tier is decided first. When the integration with a generative model already works, the temptation is to use it for everything, because adding one more question takes five minutes. That's how a product piles up calls nobody designed, and data that leaves without anyone deciding it.
- Key management is a product problem. What happens when the salon changes tablet? If the answer isn't designed as well as the rest of the interface, "the server can't read it" becomes "nobody can read it anymore".
- Tests must be prepared before they are needed. If you don't keep production data, you can't use it to try a new model. The test set has to be built in advance, and the criteria for an "acceptable result" written before the migration, not during.
And in your company?
The method is the same as in the services I offer, even when the industry changes.
- Orders arriving by email. "Is it an order? From which customer? Which lines?" are closed decisions: they get extracted and land in the ERP without retyping, without bothering a generative model for every message.
- Production photos. Image checks can sit next to the machine, as the colour try-on sits on the tablet. Photos don't necessarily have to leave the shop floor.
- ERP and e-commerce. Often the right tier is "no AI": stock levels kept in sync between ERP and shop are an integration, and an integration never has to guess.
- Machines on the shop floor. States, part counts and alarms read over OPC UA, MQTT or Modbus are structured data. Before thinking about a model, they have to be delivered where they are needed.
Intro call · 30 minutes · free
Want to know where AI belongs in your process, and where it doesn't?