Olivv Assets Classifier: Vision Pipeline
An AI vision pipeline that sorts restaurant images into a strict, machine-readable contract, so a website builder can fill its template slots from data instead of asking a human.
The problem
Olivv builds restaurant websites from templates, and templates have slots: a hero banner, menu cards, a gallery, a team section. Something has to decide which photo goes where, and doing that by hand stops scaling after the first few hundred images.
Vision models make it harder than it sounds. Left unchecked they invent categories, leave fields blank, and mislabel confidently. Anything wrong that reaches storage comes back out later, when the website builder queries it.
Who it’s for
Two very different users share one pipeline. Developers stocking the asset library in bulk, and restaurant owners uploading a dozen of their own photos and reviewing each one in a browser workbench.
How it works
Every image is read by a vision model and forced through a strict contract before anything is stored: what it shows, which template slots it can fill, alt text, and tags. An answer that doesn’t fit is rejected and asked again, rather than stored and patched later.
Nothing is resolved silently. Duplicates are refused, uncertain answers wait for human approval, and batch saves are all-or-nothing. When the builder later asks which image fits a menu card, it gets an honest verdict: a real match, a fallback that merely fits, or no image at all. It never puts a burger next to the biryani label.
Decisions worth naming
- A wrong answer costs a retry, never a bad record. Invalid answers are asked again, then escalated to a different model, so one vendor’s bad day never reaches the library.
- Two speeds, one pipeline. Bulk loading runs on the rules and flags the exceptions. Owner uploads stop for review, because it’s their restaurant and they read every line.