How to scrape JavaScript-rendered websites with Python

Scrape JavaScript-generated content with Python and Playwright. Wait for rendered elements, extract structured data, save HTML, and capture screenshots.

Blog post6 min read

Written by

Dmytro Krasun

Published on

To scrape a JavaScript-rendered website with Python, load it in a browser, wait for its content, and extract the resulting DOM. Playwright is a practical way to do that locally.

A normal HTTP request can return the page’s initial HTML before the application has created its product cards, quotes, or report rows. Parsing that response gives you the initial document. A browser executes the scripts and gives you access to the updated page.

This tutorial uses Quotes to Scrape’s JavaScript demo, a practice website. We will save structured quote data, the rendered HTML, and a screenshot. For screenshot-specific settings, see the existing Python website screenshot guide.

Choose how to retrieve the content

Before launching a browser, check where the data comes from:

Where the data is availablePractical approach
In the HTML responseDownload it with requests and parse it with Beautiful Soup.
In an available JSON endpoint or embedded JSONRead that data directly when it fits your task.
In the DOM after JavaScript runsRender the page with Playwright, then extract it.
In rendered HTML that you want without operating a browserRequest HTML from ScreenshotOne, then parse it.

A page can contain its data inside a script even when the visible cards are created later. Browser rendering is useful when you want to work with the displayed elements or capture an image of them.

Install Playwright for Python

Create a virtual environment and install Playwright and Chromium:

Terminal window
python -m venv .venv
source .venv/bin/activate
python -m pip install playwright
python -m playwright install chromium

On Windows, activate the environment with .venv\Scripts\activate instead of the source command.

Scrape the rendered page and save the results

Save this script as scrape-quotes.py:

import json
from pathlib import Path
from playwright.sync_api import sync_playwright
URL = "https://quotes.toscrape.com/js/"
with sync_playwright() as p:
browser = p.chromium.launch()
try:
page = browser.new_page(viewport={"width": 1280, "height": 800})
response = page.goto(
URL, wait_until="domcontentloaded", timeout=30_000
)
if response is not None and response.status >= 400:
raise RuntimeError(f"Page returned HTTP {response.status}")
page.locator(".quote .text").first.wait_for(
state="visible", timeout=20_000
)
quotes = []
for card in page.locator(".quote").all():
quotes.append({
"text": card.locator(".text").inner_text().strip(),
"author": card.locator(".author").inner_text().strip(),
"tags": [
tag.strip()
for tag in card.locator(".tag").all_text_contents()
],
})
if not quotes:
raise RuntimeError("No quotes found; check the page and selectors")
Path("quotes.json").write_text(
json.dumps(quotes, ensure_ascii=False, indent=2),
encoding="utf-8",
)
Path("rendered.html").write_text(page.content(), encoding="utf-8")
page.screenshot(path="quotes.png", full_page=True)
print(f"Saved {len(quotes)} quotes, rendered.html, and quotes.png")
finally:
browser.close()

Run it with:

Terminal window
python scrape-quotes.py

The output gives you three views of the same page visit:

  • quotes.json contains the fields your application can process.
  • rendered.html contains the document after the readiness check.
  • quotes.png shows what the browser displayed at capture time.

The selectors match the practice site. Replace them when working with your own page. The screenshot is particularly useful when debugging an unexpected empty result or checking whether the extracted data came from the right page state.

Wait for the data, not just navigation

The example waits for the first visible quote text. That is a useful signal for this demo, which inserts its quote list together. It does not prove that every item in a different application has finished loading.

For a report that streams rows progressively, wait for its completion marker. If you control the page, adding an attribute such as data-ready="true" makes the contract explicit:

page.locator('.report[data-ready="true"]').wait_for(
state="visible", timeout=20_000
)

This is a replacement selector for a page with that marker; it is not part of the Quotes to Scrape demo.

Calling locator.all() reads the current matching elements. Choose your readiness condition before taking that snapshot. A fixed sleep can be too short on a slow page and waste time on a fast one. Network silence also does not establish that a particular chart or list is complete.

See Playwright’s locator waiting documentation and our guide to waiting for a page to load for more examples.

Parse rendered HTML with Beautiful Soup

If you already have parsing code built around Beautiful Soup, keep it. Render the page first and pass its HTML to the parser.

Install Beautiful Soup:

Terminal window
python -m pip install beautifulsoup4

The first script saved rendered.html. You can parse that file in a separate step:

from pathlib import Path
from bs4 import BeautifulSoup
html = Path("rendered.html").read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
for card in soup.select(".quote"):
text = card.select_one(".text")
author = card.select_one(".author")
if text is not None and author is not None:
print(author.get_text(strip=True), text.get_text(strip=True))

Beautiful Soup parses the supplied HTML; it does not launch the browser or run the page’s scripts. That makes the rendering step essential when your chosen fields are missing from the initial document.

The saved HTML is a document snapshot. It does not package every external image, stylesheet, or script into a portable website archive.

Get rendered HTML through ScreenshotOne

For an application that needs rendered content without running a local browser, ScreenshotOne supports format=html. You can feed that response into the same parser.

Install the HTTP client and set your API access key:

Terminal window
python -m pip install requests beautifulsoup4
export SCREENSHOTONE_ACCESS_KEY="your_access_key"

Save this as scrape-with-api.py:

import os
from pathlib import Path
import requests
from bs4 import BeautifulSoup
response = requests.get(
"https://api.screenshotone.com/take",
headers={"X-Access-Key": os.environ["SCREENSHOTONE_ACCESS_KEY"]},
params={
"url": "https://quotes.toscrape.com/js/",
"format": "html",
"wait_for_selector": ".quote .text",
},
timeout=120,
)
response.raise_for_status()
html = response.content.decode("utf-8")
Path("rendered.html").write_text(html, encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
cards = soup.select(".quote")
if not cards:
raise RuntimeError("No quotes found; inspect rendered.html")
for card in cards:
text = card.select_one(".text")
if text is not None:
print(text.get_text(strip=True))

Use the access key from your ScreenshotOne account. Run this as a server-side script, where the key is available in the environment.

ScreenshotOne’s wait_for_selector waits for a matching element to exist in the DOM. It does not require visibility or guarantee that a whole collection is complete. Choose a selector that represents the data you need. See the rendering options.

When your application needs both content and an image, request a screenshot and page content in the same render. If your destination is a knowledge base or an LLM workflow, our website-to-Markdown tutorial shows how to request a text-oriented output instead.

Handle pagination and lazy-loaded content

The scripts above scrape one page. A full crawl needs an explicit plan for discovering and visiting more URLs.

For pagination, extract the next-page link, resolve it against page.url, and navigate to that URL. Repeat the content readiness check on every page. Set a maximum page count and track visited URLs so a repeated link does not create an endless loop.

For infinite scroll, use bounded scrolling and wait for the item count or a completion marker to change. Stop at a limit even if the page can keep loading. A full-page screenshot captures the available page height; it does not automatically discover every record in an infinite list.

Some applications virtualize their lists and keep only visible rows in the DOM. In that case, extracting the DOM once cannot give you the entire dataset. Collect rows as you scroll, use pagination, or inspect an available data endpoint.

For a Python workflow that also discovers and processes multiple pages, see the Crawl4AI screenshot guide.

Diagnose empty or incomplete results

SymptomWhat to check
The HTTP response has no matching cardsCompare the initial HTML with the rendered document. The data may be inserted by JavaScript.
Waiting times outCheck the selector, the loaded URL, the response status, and whether the page requires a login or interaction.
Only some records are savedWait for completion, paginate, or handle a virtualized list.
The screenshot shows a loading screenYour readiness condition fired before the visible state you wanted.
The saved HTML contains a different layoutSet a consistent viewport and check redirects or device-specific content.

Keep a screenshot alongside a failed extraction when possible. It often makes a selector problem or an unexpected page state much easier to understand. For image quality and capture settings, continue with the Playwright Python screenshot tutorial.

Frequently Asked Questions

If you read the article, but still have questions. Please, check the most frequently asked. And if you still have questions, feel free reach out at support@screenshotone.com.

How do I scrape a JavaScript website with Python?

Open the page in a browser with Playwright, wait for the elements that contain your data, and extract their text or attributes. You can also request rendered HTML from a service such as ScreenshotOne and parse it in Python.

Can requests and Beautiful Soup execute JavaScript?

No. requests downloads an HTTP response, and Beautiful Soup parses the HTML you give it. For content created by JavaScript, first render the page in a browser or retrieve the underlying data from an available endpoint.

Does waiting for the page load event guarantee that all data is ready?

No. JavaScript can fetch and display data after the load event. Wait for a selector or application state that represents the content you want to extract.

Read more Crawling and scraping

Interviews, tips, guides, industry best practices, and news.

View all posts

Automate website screenshots

Exhaustive documentation, ready SDKs, no-code tools, and other automation to help you render website screenshots and outsource all the boring work related to that to us.