Back to Blog

Best JavaScript Web Scraping Libraries: Choose the Right Layer

Ava Wilson
Ava Wilson

Expert in Web Scraping Technologies

29-Sep-2026

TL;DR:

  • JavaScript scraping libraries belong to different layers. HTTP retrieval, HTML parsing, browser rendering, and crawl orchestration solve separate parts of the job.
  • Cheerio and htmlparser2 work with markup. They do not execute a site's client-side application simply because the source language is JavaScript.
  • Playwright and Puppeteer control browsers. Use them when the task requires rendered state or page interaction.
  • Crawlee organizes larger collection workflows. Choose its execution layer to match the target instead of assuming every crawl needs a browser.
  • Scrapeless Agent Browser supplies remote browser infrastructure. The library still owns the extraction logic and application-level validation.

Introduction: Select the Layer Before the Library

A page's downloaded HTML and its rendered browser state can contain different data. Selecting a library before checking that difference often creates unnecessary browser work or a parser that never sees the required fields.

JavaScript web scraping libraries span several categories. A parser turns markup into a queryable structure. A browser automation library drives an actual browser. A crawl framework coordinates requests and storage around an execution engine. These tools can compose; they are not all substitutes.

This comparison uses the page representation as the starting point. If your workflow includes browser-managed downloads, the Puppeteer download-file workflow covers the additional file-handling boundary.

JavaScript Scraping Libraries at a Glance

Choose by the missing capability in your application rather than by a generic popularity ranking.

Library or Surface Layer Useful When Does Not Replace
Native fetch HTTP retrieval The response already contains the data HTML parsing or browser rendering
Cheerio Selector-based HTML parsing You want familiar queries over markup A browser engine
htmlparser2 Parsing and event callbacks You need direct control over parsed tokens Page script execution
Playwright Browser automation Rendered content and interaction are required Your data contract
Puppeteer Browser automation Browser control fits an existing workflow Crawl policy and record validation
Crawlee Crawl orchestration Queues, handlers, and storage need coordination The underlying parser or browser

Inspect the Response Before Adding a Browser

The first decision is whether the required content exists in the response body. Retrieve an authorized sample, identify the field-bearing region, and compare it with what the browser displays.

The HTML parsing model turns markup into a document structure. Script execution can later change that structure. A parser cannot infer arbitrary application behavior from an empty container and a script reference.

Sometimes the response already includes usable structured data. Inspect documented or permitted sources before choosing a visual selector. If the field is absent from the response and appears only after interaction, move to the browser layer and wait for a specific readiness condition.

Cheerio: Query HTML Without Running the Page

Cheerio loads markup and exposes a selector-oriented API that suits many static extraction tasks. Its small conceptual surface is useful when you already have the response and need titles, links, or table cells.

Use selectors tied to meaningful document structure. A heading or a stable attribute is usually easier to review than a generated styling class. Validate the number and meaning of matches rather than trusting that a selector returning something has found the right element.

Cheerio does not render the page or run its scripts. If an element exists only after an application request completes in a browser, loading the original HTML into Cheerio will not create it.

htmlparser2: Work Directly with Parsing Events

Htmlparser2 provides parsing interfaces that can process tag and text events. It fits transformations where a selector tree is unnecessary or where the application wants explicit control over a narrow slice of markup.

That control also makes your own state handling visible. If you track a title element, close the state at the matching end tag. If you collect nested text, define how child elements affect the result. The parser supplies events; the application supplies the meaning.

For conventional page extraction, compare the clarity of a Cheerio query with the code needed to manage parser events. Fewer dependencies do not automatically mean less maintenance.

Playwright and Puppeteer: Use a Browser for Rendered State

Playwright and Puppeteer automate browsers and expose page navigation, interaction, and inspection. They are appropriate when the required content depends on script execution, a selected filter, or a permitted interaction.

The browser still needs an acceptance condition. A navigation event is not proof that an asynchronously loaded product grid is ready. Prefer a specific visible element or data-bearing attribute, with a bounded wait, then validate the extracted fields.

A document's observable state follows the DOM tree and event model. Keep actions and observations associated with the same intended page and session. Reusing a browser while losing the required cookies or selected location can change the result.

Choose between these libraries based on the existing application and required browser capabilities. This guide does not assign a universal speed winner. A browser launch, target scripts, network transfer, and parsing all contribute to the measured workload.

Crawlee: Add Orchestration When the Job Requires It

Crawlee provides a framework around crawling, including request handling, queues, and storage, with HTTP and browser-based crawler options. It fits a job that has grown beyond a single fetch-and-parse operation.

Choose a crawler type that matches the page. An HTTP-oriented path is suitable when the source representation has the needed fields; a browser-oriented path is appropriate for rendered interaction. Keep the extraction function small enough to test independently of scheduling.

A framework does not define which URLs your organization is permitted to collect. Set scope, concurrency, and completion limits explicitly. Respect the Robots Exclusion Protocol where applicable, and assess authorization and terms separately.

Start Scraping with Scrapeless

Power up your web scraping and automation workflow with Scrapeless!
Sign up today and get $5 in free credit — no credit card required.

Claim your free credit now in the Scrapeless Dashboard.

Compare Two Parsers on the Same Public Document

A shared source makes parser behavior easier to inspect. The following example downloads a public specification and extracts its title with Cheerio and htmlparser2. It confirms a narrow parsing contract without claiming that either parser renders JavaScript.

Prerequisites and Install

Use a current Node.js release supported by the selected packages, with outbound HTTPS access. The code uses ES modules, so save it as parse_document.mjs. The separate cloud-browser example requires a Scrapeless API key and remains pending authenticated execution without one.

Install the parser packages:

bash Copy
npm install cheerio@1.2.0 htmlparser2@12.0.0

Run this full script with node parse_document.mjs:

javascript Copy
import * as cheerio from 'cheerio';
import { Parser } from 'htmlparser2';

const url = 'https://www.rfc-editor.org/rfc/rfc9114.html';
const response = await fetch(url, {
  signal: AbortSignal.timeout(30000),
  headers: { 'User-Agent': 'ParserComparison/1.0' }
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
const $ = cheerio.load(html);
const cheerioTitle = $('title').text().trim();
let insideTitle = false;
let parsedTitle = '';
const parser = new Parser({
  onopentag(name) { if (name === 'title') insideTitle = true; },
  ontext(text) { if (insideTitle) parsedTitle += text; },
  onclosetag(name) { if (name === 'title') insideTitle = false; }
});
parser.write(html);
parser.end();
parsedTitle = parsedTitle.trim();
if (cheerioTitle !== parsedTitle || !parsedTitle.includes('HTTP/3')) {
  throw new Error('Expected document title missing or parser mismatch');
}
console.log(JSON.stringify({
  source: response.url, title: parsedTitle, parsers_agree: true
}, null, 2));

Both paths must agree on the expected title for this document. Agreement on one page is not proof that malformed markup or every site's structure will behave identically. Retain representative source captures when the extraction contract expands.

Where Scrapeless Agent Browser Fits

Scrapeless Agent Browser supplies a remote browser session that a compatible browser library can control. The service is infrastructure; Cheerio, Puppeteer, or another client still determines what the application extracts and accepts.

The Node.js SDK browser example prepares browserWSEndpoint, which Puppeteer uses to connect. Keep that value private: connection URLs can carry session authority and should not appear in application logs.

Install the SDK and browser client into the same project:

bash Copy
npm install @scrapeless-ai/sdk@1.12.1 puppeteer-core@25.12.0

Note: This cloud-browser example requires SCRAPELESS_API_KEY and an enabled browser service. Package imports and the SDK interface are checked; the remote browser connection and page retrieval remain pending live verification without credentials.

javascript Copy
import { Scrapeless } from '@scrapeless-ai/sdk';
import puppeteer from 'puppeteer-core';

const client = new Scrapeless({
  apiKey: process.env.SCRAPELESS_API_KEY
});
const { browserWSEndpoint } = await client.browser.create({
  sessionName: 'public-document-check',
  sessionTTL: 180,
  proxyCountry: 'US'
});
const browser = await puppeteer.connect({ browserWSEndpoint });
try {
  const page = await browser.newPage();
  await page.goto('https://www.rfc-editor.org/rfc/rfc9114.html', {
    waitUntil: 'domcontentloaded', timeout: 30000
  });
  const title = await page.title();
  if (!title.includes('HTTP/3')) throw new Error('Unexpected document');
  console.log(JSON.stringify({ title, content_check: 'passed' }));
} finally {
  await browser.close();
}

The public document is intentionally simple: it tests the connection and a content assertion before adding application-specific interaction. A production page requires its own readiness and extraction conditions. The session lifetime is a selected example setting, not a promise about every task's duration.

Review Scrapeless pricing separately from library selection. Running a parser locally and provisioning a remote browser consume different resources, even if both produce a JSON record.

Choose by the Missing Capability

A practical decision sequence starts with source availability and ends with workflow ownership.

Observation Next Choice Acceptance Check
Data exists in response HTML Fetch plus a parser Correct fields and stable identity
Data appears only after scripts run Browser library Specific ready state and required content
Task requires approved clicks or filters Browser workflow State after each action
Many URLs need lifecycle management Crawl framework Scope, deduplication, and completion rules
Browser hosting is the operational burden Managed browser infrastructure Authenticated connection and session cleanup

Do not add all layers by default. Each layer introduces a lifecycle to manage and a boundary where incomplete data can be mistaken for a finished result.

Conclusion: Keep Extraction Separate from Execution

Inspect the representation, select the smallest suitable execution layer, and validate the extracted record. Use Cheerio or htmlparser2 for markup, a browser library for rendered interaction, and Crawlee when orchestration is needed. Add remote browser infrastructure when hosting is the problem, while keeping the application's field checks explicit.

Ready to Build Your Web Data Workflow?

Join developers discussing practical collection workflows: Discord · Telegram.

Create an account at app.scrapeless.com and start with an authorized task whose output you can validate.

FAQ

Q: Does Cheerio execute JavaScript?

Cheerio parses and queries markup; it does not run the site's browser application. Content created only by page scripts requires another retrieval or rendering path.

Q: Are Playwright and Puppeteer HTML parsers?

They are browser automation libraries. They can inspect a browser's page state, but their role differs from a standalone parser that operates on supplied markup.

Q: When should a scraper use Crawlee?

Use Crawlee when request queues, handlers, and storage coordination justify a crawl framework. A single-page extraction may be simpler without that layer.

Q: Is Agent Browser a replacement for a JavaScript library?

Agent Browser supplies remote browser infrastructure. Your library and application still control page interaction, extraction, and record validation.

Q: Can a parser and a browser be used together?

Yes. A browser can obtain rendered markup that a parser later processes, provided the application preserves source context and checks that the required state was reached.

At Scrapeless, we only access publicly available data while strictly complying with applicable laws, regulations, and website privacy policies. The content in this blog is for demonstration purposes only and does not involve any illegal or infringing activities. We make no guarantees and disclaim all liability for the use of information from this blog or third-party links. Before engaging in any scraping activities, consult your legal advisor and review the target website's terms of service or obtain the necessary permissions.

Most Popular Articles

Catalogue