Product titles and images from bol.com, without scraping a single page
Where Depotely gets a product's title, brand, images and bol.com link after an import, why it never reads the product page, and how a 4,000-product catalogue is filled in the right order.
· 3 min read
An offer export from bol.com tells you what you sell and for how much. It does not tell you what the product is called or what it looks like. Depotely fills that in after every import, and the way it does so changed completely after one lesson: bol.com answers a web page request from a server with a refusal. The scrape that used to carry every title returned nothing in production, 3,819 products stayed nameless, and every blocked batch burned twelve minutes in retries. Nothing in Depotely reads a product page any more.
bol.com's own catalogue, through your own credentials
Content comes from bol.com's Retailer API, with the same credentials you connected the channel with:
- The catalogue content call, by EAN: the title, the description, the brand and bol.com's own SEO slug.
- The assets call, by EAN: the product images.
- The product id lookup: bol.com's internal id for the product, needed only to build a link.
Everything is keyed on the EAN, which the offer export already has. Nothing waits for an id that still has to be looked up.
Filled in stages, highest stock first
After an import, the EANs that still miss content go into a queue in the database. Workers take 45 rows at a time and finish one stage before starting the next: every title first, then every set of images, then the ids. The reason is bol.com's rate limits. Titles can be fetched at eight per second, images at fifty per minute. Asking for images in between would make every title wait for an image.
Within a stage the order is highest stock first, so the products that sell get their name and picture before the long tail. The inventory page can also raise the rows you are looking at to the front of the queue; the last page you looked at wins.
"Fetch details now"
One row at a time there is a button on the inventory page. It runs all three stages for that product immediately, on its own lane, so it never waits behind a running import. If the product was still queued in that import, the queued row is marked as done by hand and skipped.
Rate limits that are shared
bol.com's limits are per seller account, not per server. Depotely keeps one shared counter per endpoint and per seller, so four workers together never send more than one seller is allowed. Two imports started an hour apart do not open two queues over the same products either: an EAN that is complete, or still queued, is skipped.
Interrupted work resumes
A deploy in the middle of a 4,000-product import used to leave the job stranded. Now a batch that fails part-way puts its rows back, and a job whose queue has not moved for fifteen minutes is picked up again on its own. You see a running job on the inventory page until the queue is truly empty.
The link to the product page
A bol.com product link is built from bol.com's product id and slug, never from the EAN: the EAN URL is a dead page. While the id is still unresolved, the link opens bol.com's search for the EAN instead, so there is never a broken link in the table.
Two more things from bol.com
- "Not for sale" reasons. Not paused does not mean for sale: bol.com can refuse to publish an offer for a price that is too high, missing stock or missing selling rights. The reasons come from bol.com's unpublished-offers report and appear as a badge on the row, in bol.com's own words.
- Your shop logo. The retailer API has no logo field, so Depotely reads it once from your public seller page and keeps only an image hosted by bol.com. One attempt, no retries; a blocked page simply means no logo.
What this means for you
Connect, import, and come back in a while: titles arrive within minutes for a large catalogue, images follow, and the table fills from the products that matter most. No page is scraped, no credential leaves your account, and nothing is asked of bol.com twice.