Guide, measured 2026-10-05
How to scrape LinkedIn Jobs and Indeed without an account
Both boards show their job listings to anyone, signed in or not. Reading them in bulk is two very different problems: LinkedIn answers a plain HTTP request and stops at about 1,000 jobs per search, while Indeed answers 403 to anything that is not a real browser on a real connection. Here is what each board actually serves a logged-out visitor, and how to read past the limits.
What a logged-out visitor can see
A job listing is published to be found, and both boards let search engines and visitors read it without an account. That is the data a LinkedIn jobs scraper or an Indeed scraper works with: title, company, location, posting date, salary when the employer gives one, the full description, and the links. What a visitor does not see stays out of reach without an account, and an honest scraper does not reach for one: profiles, recruiter contact details, and on LinkedIn the employer's own apply link.
LinkedIn Jobs: two guest endpoints, plain HTTP
LinkedIn's public job search page, as served to a visitor, fills itself from two endpoints under /jobs-guest/jobs/api/. The search returns 10 job cards of HTML per request; the job page returns the body of one job. Both answer a plain HTTP client with no cookie, from a laptop and from Apify's datacenter proxy.
| LinkedIn: what we asked | What came back |
|---|---|
| Guest search, plain HTTP | 200, 10 job cards per request, from a laptop and from Apify's datacenter |
| Guest search, start=1000 | 400: one search shows about 1,000 jobs |
| Guest job page | 200: description, seniority, employment type, industries, applicants, base pay when published |
| Apply button | Sign-in page: no employer URL for visitors |
| Job type, workplace, experience and sort parameters | Ignored by the guest search (same job IDs with and without them) |
The guest search ignores the job type, workplace, experience and sort parameters that the signed-in site uses: we compared the job IDs returned with and without each one. So those filters have to be applied after reading each job's page, which is what seniority and employment type are on.
This reads the first page of a search and one job, in about 30 lines and no library:
// Node 22, no dependencies. LinkedIn's guest job search, logged out.
const base = "https://www.linkedin.com/jobs-guest/jobs/api";
const headers = { "user-agent": "Mozilla/5.0 (job-research script)" };
const text = (s) => s.replace(/<[^>]+>/g, "").replace(/\s+/g, " ").trim();
const search = new URL(`${base}/seeMoreJobPostings/search`);
search.search = new URLSearchParams({ keywords: "data engineer", location: "Paris, France", start: "0" });
const html = await (await fetch(search, { headers })).text();
const cards = html.split("<li>").slice(1).map((li) => ({
id: li.match(/urn:li:jobPosting:(\d+)/)?.[1],
title: text(li.match(/base-search-card__title">([\s\S]*?)<\/h3>/)?.[1] ?? ""),
company: text(li.match(/base-search-card__subtitle">([\s\S]*?)<\/h4>/)?.[1] ?? ""),
location: text(li.match(/job-search-card__location">([\s\S]*?)<\/span>/)?.[1] ?? ""),
postedAt: li.match(/datetime="([\d-]+)"/)?.[1],
}));
console.log(cards.length, "cards");
console.log(cards.slice(0, 3));
// One job page: description and the "job criteria" list.
const job = await (await fetch(`${base}/jobPosting/${cards[0].id}`, { headers })).text();
const criteria = [...job.matchAll(/description__job-criteria-subheader">([\s\S]*?)<\/h3>[\s\S]*?description__job-criteria-text[^>]*>([\s\S]*?)<\/span>/g)]
.map((m) => [text(m[1]), text(m[2])]);
const description = text(job.match(/show-more-less-html__markup[^>]*>([\s\S]*?)<\/div>/)?.[1] ?? "");
console.log(Object.fromEntries(criteria));
console.log(description.length, "characters of description");10 cards
[
{
id: '4446387035',
title: 'Senior Data Engineer - BeReal',
company: 'BeReal.',
location: 'Paris, Île-de-France, France',
postedAt: '2026-10-01'
},
{
id: '4440458716',
title: 'Data Engineer Expérimenté F/H',
company: 'EY',
location: 'Paris, Île-de-France, France',
postedAt: '2026-09-22'
},
{
id: '4453366502',
title: 'Senior Data Engineer - Central Data & Analytics Team',
company: 'Voodoo',
location: 'Paris, Île-de-France, France',
postedAt: '2026-09-24'
}
]
{
'Seniority level': 'Not Applicable',
'Employment type': 'Full-time',
Industries: 'Software Development'
}
2683 characters of descriptionTwo details the output shows. Titles come HTML-encoded (&), so decode entities before storing them. And the card date is the employer's posting date, not the time you read it.
What LinkedIn does not show a visitor
- The employer's apply URL.The Apply button opens "Join or sign in to find your next job", and neither the guest page nor the full job page carries the URL. We read 4,440 rows from Apify's cloud through datacenter and residential addresses in the US and France: 0 had it. The page does say whether the job uses Easy Apply or the employer's site, which is worth keeping.
- A workplace field.Remote, hybrid and on-site are a filter for signed-in members only. The title, the location and the description sometimes say it ("Remote", "hybrid, 2 days on site"): 0 to 13% of rows in our runs.
- Salary, most of the time. The job page shows base pay when the employer publishes it: 5 to 15% of rows in Paris and London, the same share as the paid LinkedIn scrapers we measured on the same searches (6 to 14%).
Going past 1,000 LinkedIn jobs
The guest search takes start in steps of 10 and answers 400 from start=1000, so one search shows about 1,000 jobs, a few of them repeated across pages. The date filter (f_TPR) is one of the few parameters the guest search honours, so the same search in narrower windows (past month, past two weeks...) reaches jobs the first 1,000 did not. Union them by job ID and stop when nothing new comes back. On a London nurse search, the first window gave 970 unique jobs, the past month brought it to 1,143 and two weeks to 1,299.
Pace matters more than proxies here. Our first cloud build sent every request through one datacenter address and drew enough HTTP 429 answers to take 14 to 18 minutes per 1,000 jobs with their pages. Spreading the same requests over 16 proxy sessions cut it to about 5.5 minutes, with no residential traffic at all.
Indeed: a 403 before anything else
Indeed's search page sits behind Cloudflare. A plain HTTP request gets a 403 challenge page, re-checked with curl and a current Chrome User-Agent on 2026-10-05. After that, it depends on who is asking:
| Indeed: what we asked | What came back |
|---|---|
| Search page, plain HTTP | 403, Cloudflare challenge |
| Search page, headless Chromium as shipped | "Blocked": its User-Agent says HeadlessChrome |
| Search page, headless Chromium with a browser User-Agent, home connection | 200 |
| Search page, Apify cloud without proxy | Connection refused |
| Search page, Apify datacenter proxy | 403, "Security Check" |
| Search page, Apify residential proxy in the country | 200, 15 jobs on the first page |
| Page 2 and further, the job page | Challenge, then a sign-in page or HTTP 401 |
Three consequences. You need a real browser, with a browser's User-Agent. In the cloud you need residential proxy in the searched country, which makes Indeed the expensive board of the two even though its pages are lighter. And you cannot walk page after page or open each job the way you would on LinkedIn: a scraper that does gets a sign-in page within a few requests, and our own address was challenged after about 600 page loads in a day.
Our Actor therefore reads Indeed from inside one real browser session on its public search page, through residential proxy, with the full descriptions, 100 jobs per request; when that path is refused, it falls back to reading search pages one by one. Indeed's search ends at 1,000 jobs, so past that it reads narrower searches (by posting date and job type) and unions them by job key: 1,812 jobs for a Paris search Indeed counts at 1,821.
What it costs in the cloud
Measured on Apify with our Actor: the run's total usage, compute and proxy included, for 1,000 jobs.
| Run | Jobs | Run time | Apify usage |
|---|---|---|---|
| LinkedIn, data engineer, Paris, with job pages, datacenter proxy | 1,000 | 329 s | $0.025 |
| LinkedIn, nurse, London, with job pages, datacenter proxy | 1,000 | 334 s | $0.025 |
| Indeed, nurse, London, with descriptions, residential proxy | 1,000 | 45 s | $0.029 |
| Indeed, data engineer, Paris, with descriptions, residential proxy | 1,000 | 63 s | $0.063 |
| Indeed, data engineer, Paris, past 1,000, residential proxy | 1,812 | 136 s | $0.166 |
LinkedIn costs compute, not bandwidth: about 1,180 requests per 1,000 jobs with their pages, on datacenter proxy. Indeed costs residential traffic, and a search past 1,000 rereads overlapping searches, so its cost per job rises with depth.
The traps that make a job dataset wrong
- A city is an area.Both boards widen "London" to its surroundings (Teddington, Morden). If you need the city itself, filter on the location after reading.
- Indeed's search is broad.It matches the keywords anywhere in the job: 7 of the first 20 Indeed titles for "data engineer" in Paris named data, against 20 of 20 on LinkedIn.
- The date you store.Indeed exposes both the employer's publication date and the date Indeed indexed the job; one paid Indeed scraper we measured returns no posting date at all, another returns the indexing date.
- One-sided salaries."From £37.50 an hour" has a minimum and no maximum. A parser that expects a range drops it: on the same 100 London jobs, one Indeed scraper filled salary on 58 rows and another on 90.
- Success is not data.A run can finish "succeeded" while a board returned nothing. Count rows per board after every run, and report a short result as partial with the reason.
Or skip the plumbing
Our LinkedIn Jobs and Indeed scraperdoes all of the above in one run: both boards, one row format, salary parsed, the employer's posting date, the full description, past 1,000 jobs, and a partial status instead of a silent short result. It costs $0.35 per 1,000 LinkedIn jobs and $0.10per 1,000 Indeed jobs, descriptions included, on the Apify Store. For the employers' own career pages (Greenhouse, Lever, Ashby and four more), with the direct apply link LinkedIn hides, see JobsRadar and the public ATS job board APIs.
Questions
Can you scrape LinkedIn Jobs without logging in?
Yes. LinkedIn's public job search and job pages are served to logged-out visitors, and the two endpoints those pages call answer a plain HTTP request: 10 job cards per request, and each job's description, seniority, employment type, industries and applicant count. No account, no cookie.
How many results does LinkedIn show per job search?
About 1,000. The guest search takes an offset in steps of 10 and answers HTTP 400 from start=1000 on. Past that, the same search split into date windows (past day, week, month) returns further jobs: 1,299 unique cards for a London nurse search where the first window gave 970.
Can you get the apply URL from LinkedIn without an account?
No. A logged-out visitor who clicks Apply gets a sign-in page, and neither the guest job page nor the search card carries the employer's URL. We found it on 0 of 4,440 rows read from Apify's cloud through datacenter and residential addresses. What the page does say is whether the job uses Easy Apply or the employer's site.
Why does Indeed return 403?
Indeed's search sits behind Cloudflare, which answers a plain HTTP client with a 403 challenge page (re-checked 2026-10-05). A real browser passes from a home connection; from cloud addresses Indeed refuses or challenges it, and through residential proxy in the searched country it serves the page.
Does reading the public pages need a browser?
For LinkedIn, no: the guest endpoints answer plain HTTP, and a Node script with fetch reads them (the sample above). For Indeed, yes: a plain request gets 403, and a headless browser must not announce itself as HeadlessChrome.
Public job listings only: no login, no account, no profiles. Not affiliated with LinkedIn or Indeed.