Engineering
Notes from the team building Favora: the crawlers, the catalogue, and the models that make fashion discovery personal. Real systems, live numbers.
Live catalogue Latest catalogue activity 2 Oct, 13:01 CEST
2,958,140 product images in the catalogue
9,484,242 pages crawled in the last 30 days
514,605 items catalogued, one schema
187 brands crawled, enriched and indexed
Writing
The Favora dataset
How we turn the fashion internet into one machine-readable catalogue, why we build it ourselves, and the current scale of the dataset.
Read the piece
The stack
Python and Rust pipeline workers on Temporal. Postgres for the catalogue, Typesense for keyword, faceted and vector search. An LLM labels and embeds every item, gated on a hand-labelled gold set. SvelteKit on the web, Expo in the app.
Working on consumer AI, search, or crawling at scale? We are hiring. For everything else: humans@favora.ai