News · New this week
Semantic photo search with CLIP in the browser, built in 29 minutes
Day 1 of a vibe coding challenge: one spin, one hour, no rerolls. The result runs CLIP in the browser with no backend and no API key.
The rules were one spin, one hour, no rerolls. A random spinner picked the AI app to build, and it landed on semantic photo search. The app was built and deployed live on Vercel at 29:12 of the 60-minute clock.
Here's what was built and why it works without a server.
Research first, then the prompt
Before writing anything, I asked Claude on the web how photo search even works. The answer was CLIP, an open model that understands both photos and words.
The specific model is CLIP ViT-B/32, run through Transformers.js. That choice shaped the whole build, because it means the model can live in the browser.
Then I typed the build prompt into Claude Code by hand. Claude Code built a Vite + React app from it. After the prompt, my job was mostly watching.
How the search works
CLIP puts photos and text into the same kind of space, so you can compare them directly. Every photo goes through CLIP and becomes 512 numbers. Those numbers are cached in IndexedDB, so they're saved right in your browser.
When you search, your text becomes 512 numbers too. The app compares them with the photo numbers, and the photo with the highest similarity score wins.
The scores show the idea. For a photo of cats, the text "cats lying on a couch" scored 0.31. The text "a dog on a beach" scored 0.12. The match isn't a keyword hit. It's just the higher number.
In the demo, searching "two cats sleeping on a couch" found the right photo. So did "a teddy bear". You describe the picture and the app finds it.
Where it runs
All of it happens in the browser. It runs on WebGPU with a WASM fallback. There's no backend and no API key.
That's a practical win for a one-hour build. There's no server to set up, no key to manage, and no service to call per search. Embedding the photos once and caching the results means later searches only need to embed your text.
What to try
If you want to try this yourself, start with the same two steps. Ask Claude how the feature works before you prompt for code, then write the build prompt yourself. CLIP through Transformers.js is a good fit when you want search by description with no backend. The challenge is also open to you, since the next app is still up for grabs. This was Day 1, and the question for the next spin is which app the spinner should land on.



