Web scraping with Node.js and Cheerio: static and JavaScript-rendered pages

Learn web scraping with Node.js and Cheerio. Extract structured data from static HTML, render JavaScript pages with Playwright or an API, and save JSON.

Blog post6 min read

Written by

Dmytro Krasun

Published on

For web scraping with Node.js, start by fetching the HTML and parsing it with Cheerio. When the page creates its content with JavaScript, render it first and give Cheerio the resulting HTML.

Cheerio provides familiar selectors and methods such as .find(), .text(), and .attr(). It works well for extracting a title, product cards, links, or table rows from a document. It does not run a browser, execute scripts, or create a screenshot.

This guide uses the static and JavaScript versions of Quotes to Scrape. We will reuse one parser with three ways of obtaining HTML: a normal HTTP request, ScreenshotOne, and a local Playwright browser.

Install Cheerio

Use Node.js 22.19 or later, as required by the current Cheerio installation documentation. The examples use Node’s built-in fetch and .mjs files for ES modules.

Create a project and install Cheerio:

Terminal window
mkdir cheerio-scraper
cd cheerio-scraper
npm init -y
npm install cheerio

Write a reusable extraction function

Save this as parse-quotes.mjs:

import * as cheerio from "cheerio";
export function parseQuotes(html) {
const $ = cheerio.load(html);
return $(".quote")
.toArray()
.map((element) => {
const card = $(element);
return {
text: card.find(".text").text().trim(),
author: card.find(".author").text().trim(),
tags: card.find(".tag")
.toArray()
.map((tag) => $(tag).text().trim()),
};
});
}

The parser takes an HTML string and returns plain JavaScript objects. It does not care how the HTML was obtained. Keeping extraction separate from retrieval lets you change the rendering approach without rewriting every selector.

The selectors match this practice site. For another website, inspect its document and substitute the elements that contain your own fields.

Scrape a static page with fetch and Cheerio

Save this as scrape-static.mjs in the same directory:

import { writeFile } from "node:fs/promises";
import { parseQuotes } from "./parse-quotes.mjs";
const response = await fetch("https://quotes.toscrape.com/", {
signal: AbortSignal.timeout(30_000),
});
if (!response.ok) {
throw new Error(`Page returned HTTP ${response.status}`);
}
const html = await response.text();
const quotes = parseQuotes(html);
if (quotes.length === 0) {
throw new Error("No quotes found; inspect the response and selectors");
}
await writeFile("quotes.json", JSON.stringify(quotes, null, 2), "utf8");
console.log(`Saved ${quotes.length} quotes`);

Run it with:

Terminal window
node scrape-static.mjs

The static practice page includes its quote cards in the HTML response, so this approach can extract them without launching a browser.

The examples work with UTF-8 HTML. If you are reading pages with other encodings, Cheerio also has loaders for raw bytes and streams. See its document loading documentation.

Why the same scraper can miss JavaScript content

Change the URL in the static script to https://quotes.toscrape.com/js/. That version uses JavaScript to construct its quote cards. The initial document can include scripts and data without containing the .quote elements our parser expects.

Cheerio parses the supplied document. Installing another HTML parser or waiting a few seconds after fetch will not execute the scripts inside that response.

You have several possible paths:

  • Retrieve an available data endpoint or embedded JSON when those fields are sufficient.
  • Render the page through an API and parse its HTML.
  • Use a local browser when you need navigation, clicks, or a custom browser session.

This is a retrieval decision. The parseQuotes function can stay the same once the cards exist in the HTML.

Scrape JavaScript-rendered HTML with ScreenshotOne

ScreenshotOne can render the target page and return its HTML with format=html. Your Node.js application handles the request and Cheerio handles the extraction.

Set your access key from a ScreenshotOne account:

Terminal window
export SCREENSHOTONE_ACCESS_KEY="your_access_key"

Save this as scrape-rendered.mjs:

import { writeFile } from "node:fs/promises";
import { parseQuotes } from "./parse-quotes.mjs";
const accessKey = process.env.SCREENSHOTONE_ACCESS_KEY;
if (!accessKey) {
throw new Error("Set SCREENSHOTONE_ACCESS_KEY before running this script");
}
const requestUrl = new URL("https://api.screenshotone.com/take");
requestUrl.search = new URLSearchParams({
url: "https://quotes.toscrape.com/js/",
format: "html",
wait_for_selector: ".quote .text",
}).toString();
const response = await fetch(requestUrl, {
headers: { "X-Access-Key": accessKey },
signal: AbortSignal.timeout(120_000),
});
if (!response.ok) {
throw new Error(`Rendering request returned HTTP ${response.status}`);
}
const html = await response.text();
const quotes = parseQuotes(html);
await writeFile("rendered.html", html, "utf8");
if (quotes.length === 0) {
throw new Error("No quotes found; inspect rendered.html and the selectors");
}
await writeFile("quotes.json", JSON.stringify(quotes, null, 2), "utf8");
console.log(`Saved ${quotes.length} quotes from rendered HTML`);

Run it with node scrape-rendered.mjs. Keep the access key in the server-side environment.

The wait_for_selector option waits for a matching element to appear in the DOM. For an application that inserts data gradually, use a completion marker that represents the finished content. A first card appearing does not establish that every card has loaded. See the selector wait options.

If you also need an image of the page, the API can capture a screenshot and return content in the same request. Choose HTML metadata for your Cheerio parser, or Markdown metadata for a reading workflow.

Render locally with Playwright, then parse with Cheerio

A local browser is useful when you need to click a button, log in through a custom flow, or inspect the page while developing your scraper.

Install Playwright and Chromium:

Terminal window
npm install playwright
npx playwright install chromium

Save this as scrape-browser.mjs:

import { writeFile } from "node:fs/promises";
import { chromium } from "playwright";
import { parseQuotes } from "./parse-quotes.mjs";
const browser = await chromium.launch();
try {
const page = await browser.newPage({
viewport: { width: 1280, height: 800 },
});
const response = await page.goto("https://quotes.toscrape.com/js/", {
waitUntil: "domcontentloaded",
timeout: 30_000,
});
if (response && response.status() >= 400) {
throw new Error(`Page returned HTTP ${response.status()}`);
}
await page.locator(".quote .text").first().waitFor({
state: "visible",
timeout: 20_000,
});
const html = await page.content();
const quotes = parseQuotes(html);
if (quotes.length === 0) {
throw new Error("No quotes found; check the page and selectors");
}
await writeFile("quotes.json", JSON.stringify(quotes, null, 2), "utf8");
await writeFile("rendered.html", html, "utf8");
await page.screenshot({ path: "quotes.png", fullPage: true });
console.log(`Saved ${quotes.length} quotes and a screenshot`);
} finally {
await browser.close();
}

Run it with node scrape-browser.mjs. This produces the same structured fields as the other scripts, plus a screenshot for checking the rendered state.

Playwright supplies the browser and Cheerio supplies the parser. You could also extract directly with Playwright locators; Cheerio is useful when you already have parsing functions you want to reuse. Our Playwright waiting guide explains how to pick stronger readiness conditions.

For pagination, look for the next-page link in the HTML after each extraction. Resolve relative links against the page URL with new URL(href, pageUrl). Keep a set of visited URLs and a maximum page count.

Do not assume that all pages have the same loading behavior. A static category page can link to JavaScript-heavy detail pages. You can use ordinary requests for the first group and render only the pages that need it.

When increasing throughput, bound the number of concurrent requests, give each one a timeout, and record failures with their target URL. Start with a small sample before launching a large crawl. ScreenshotOne’s concurrency documentation explains the API-side limit.

Choose HTML, Markdown, or screenshots for the result

Your application needsUseful output
Specific fields or attributesHTML parsed by Cheerio into structured objects.
Readable page content for notes, search, or an LLMMarkdown.
A record of the page’s visual stateA screenshot.
Both readable content and visual evidenceA screenshot with content metadata.

For Markdown output, continue with converting a website to Markdown with JavaScript or Python. For screenshots without extraction, use the existing Node.js website screenshot guide.

Frequently Asked Questions

If you read the article, but still have questions. Please, check the most frequently asked. And if you still have questions, feel free reach out at support@screenshotone.com.

Can Cheerio scrape a website that uses JavaScript?

Cheerio can parse HTML produced by a JavaScript application, but it cannot execute the page's scripts. First render the page with a browser or rendering API, then pass the resulting HTML to Cheerio.

What is the difference between Cheerio and Playwright?

Cheerio parses HTML and provides methods for selecting and extracting elements. Playwright controls a browser that can load pages, execute JavaScript, interact with elements, and take screenshots. They can be used together.

Do I need a browser for every Node.js scraper?

No. If the data is in the HTTP response, fetch the HTML and parse it with Cheerio. Use rendering when the fields you need are created by client-side JavaScript or depend on browser interaction.

Read more Crawling and scraping

Interviews, tips, guides, industry best practices, and news.

View all posts

Automate website screenshots

Exhaustive documentation, ready SDKs, no-code tools, and other automation to help you render website screenshots and outsource all the boring work related to that to us.