Guide, measured 2026-10-01
How to scrape a Shopify store
Every Shopify storefront publishes its whole catalog as JSON, to anyone, without a key. The hard part is not getting the data: it is knowing which endpoint has which field, how pagination ends, and which stores quietly give you nothing.
The five public endpoints
All of them answer on the store's own domain, logged out, the same way they answer a shopper's browser.
| Endpoint | What it gives | What it does not |
|---|---|---|
| /products.json?limit=250&page=N | Every product with variants (price, compare-at price, available, SKU, grams), images (size, variant links), options, tags, dates, body HTML | Stock counts, barcodes, image alt text |
| /collections/{handle}/products.json | The same, for one collection | Same as above |
| /products/{handle}.js | One product: inventory_quantity and inventory policy, barcode, media with alt text, selling plans | Dates per variant. And prices are in cents here |
| /meta.json | Store name, currency, published_products_count, published_collections_count | Anything per product |
| /collections.json | Every collection with its handle and products_count | The products themselves |
Reading the whole catalog
products.json returns at most 250 products per request: asking for limit=500 still returns 250. Page with ?page=N until a page comes back empty. Past the last page you get {"products": []} with a 200, not an error: on Allbirds, page 4 is empty because 692 products fit in three pages.
const store = "https://www.allbirds.com";
const products = [];
for (let page = 1; ; page++) {
const res = await fetch(`${store}/products.json?limit=250&page=${page}`);
if (!res.ok) throw new Error(`HTTP ${res.status} on page ${page}`);
const { products: batch } = await res.json();
if (batch.length === 0) break; // past the last page: an empty list, not an error
products.push(...batch);
await new Promise((resolve) => setTimeout(resolve, 400)); // be polite
}
console.log(`${products.length} products`);It printed 692 products, exactly the published_products_count that /meta.json announces for the store.
How long it takes
| Store | Products on page 1 | Size | Median |
|---|---|---|---|
| allbirds.com | 250 of 692 | 1.7 MB | 1.31 s |
| kith.com | 250 of 25,001 | 1.4 MB | 0.93 s |
| deathwishcoffee.com | 147 of 147 | 629 KB | 0.83 s |
The per-product Ajax endpoint answered in a median 0.49 s. One request per product is what makes stock and barcodes the slow part of any Shopify scraper: 300 products with a polite pause take about a minute and a half.
The traps
- Cents in one place, decimals in the other. The same Allbirds variant is
"25.00"inproducts.jsonand2500in/products/{handle}.js. Pick one source per field and convert. - No stock, no barcode in products.json. Both live only in the Ajax endpoint, and
inventory_quantityis only meaningful wheninventory_managementis set. Allbirds tracks stock on every variant; Kith publishes no counts at all, though it does publish a barcode on every variant. - The placeholder option. A product with a single variant still has an option called
Titlewith the valueDefault Title. It is not a real option; drop it before you show variants to anyone. - Compare prices as integers. Comparing
19.99and19.98as floats to detect a one-cent price drop fails in more than one language. Store cents. - An empty catalog is not a broken scraper. tattly.com answers
products.jsonwith an empty list, and its/meta.jsonsayspublished_products_count: 0. Check the count before you debug. - Headless stores serve nothing. gymshark.com runs Shopify behind a Next.js storefront:
/products.jsonanswers 403 and/meta.json404, whatever the user agent. There is no public catalog to read there. - Big catalogs are big. Kith announces 25,001 products: 101 requests of about 1.4 MB each. Set a cap, and remember that a capped read cannot tell you a product was removed, only that you did not reach it.
Or skip the plumbing
Our Shopify and WooCommerce scraper does all of the above, reads WooCommerce stores through their public Store API in the same run, returns 62 normalized fields per product, and tracks price and stock changes between runs. It runs on the Apify Store at $1.50 per 1,000 products.
Questions
Is products.json available on every Shopify store?
On every store that uses Shopify's own online store. It is part of the storefront, not an app. It is missing on headless storefronts, where Shopify only runs the back office: gymshark.com, a Next.js site, answers 403 to /products.json.
How many products does products.json return per request?
At most 250. Asking for limit=500 still returns 250, so page through with ?limit=250&page=N until a page comes back empty.
How do I get stock levels and barcodes from a Shopify store?
Not from products.json, which carries neither. The per-product Ajax endpoint, /products/{handle}.js, has inventory_quantity (only when the store tracks inventory), the inventory policy and the barcode, at one request per product.
Why are prices 100 times too big?
You are reading the Ajax endpoint. products.json gives prices as decimal strings ("25.00"); /products/{handle}.js gives integer cents (2500). Mixing the two is the classic bug.
How do I know how many products a store has?
/meta.json returns published_products_count, along with the currency and the number of published collections. Kith announces 25,001 products, which is 101 pages of products.json.