Convert a website URL to PDF with Node.js and Python

Save a website as a PDF with Node.js or Python. Choose paper size, margins, print or screen styles, and handle rendering errors before saving the file.

Blog post5 min read

Written by

Dmytro Krasun

Published on

To turn a website URL into a PDF, ask ScreenshotOne for format=pdf and save the response bytes to a file. The page is rendered in a real browser, including content added by JavaScript. For pages that load data later, you may need to wait for a selector or another readiness signal before rendering the PDF.

The scripts below produce an A4 document with margins, which suits most reports and archives. If you need something closer to what you see on screen, or one long page instead of several sheets, see Choose the PDF layout.

Need just one file? The URL-to-PDF tool is quicker. The code is for when PDFs are part of your product: exports, reports, or archiving. More on that on the PDF API page.

Before you start

  • Node.js 22+ or Python 3.10+. Both scripts use only the standard library, so there’s no browser or SDK to install.
  • A ScreenshotOne access key in the SCREENSHOTONE_ACCESS_KEY environment variable. The scripts send it in the X-Access-Key header. You don’t need the signing secret for server-side requests like these unless you’ve turned on mandatory signing in your account.
  • A URL that opens without logging in. The rendering browser doesn’t have your cookies.

Node.js

Save this as url-to-pdf.mjs:

import { writeFile } from "node:fs/promises";
const accessKey = process.env.SCREENSHOTONE_ACCESS_KEY;
const [inputUrl, outputPath = "website.pdf"] = process.argv.slice(2);
if (!accessKey || !inputUrl) {
throw new Error(
"Set SCREENSHOTONE_ACCESS_KEY, then run: node url-to-pdf.mjs URL [output.pdf]"
);
}
const response = await fetch("https://api.screenshotone.com/take", {
method: "POST",
headers: {
"Content-Type": "application/json",
"X-Access-Key": accessKey,
},
body: JSON.stringify({
url: new URL(inputUrl).href,
format: "pdf",
pdf_paper_format: "a4",
pdf_margin: "12mm",
pdf_print_background: true,
media_type: "print",
}),
signal: AbortSignal.timeout(120_000),
});
if (!response.ok) {
throw new Error(`PDF request failed: HTTP ${response.status}`);
}
const pdf = Buffer.from(await response.arrayBuffer());
if (pdf.subarray(0, 5).toString() !== "%PDF-") {
throw new Error("The response isn't a PDF; nothing was written.");
}
await writeFile(outputPath, pdf);
console.log(`Saved ${outputPath} (${pdf.length} bytes)`);

Run it:

Terminal window
node url-to-pdf.mjs 'https://example.com/' example.pdf

The response body is the PDF, so there’s no JSON or base64 to decode. The %PDF- check is cheap insurance: if something goes wrong, you get an error instead of a .pdf file that won’t open.

Python

Save this as url_to_pdf.py:

import json
import os
import sys
from pathlib import Path
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
def main():
access_key = os.environ.get("SCREENSHOTONE_ACCESS_KEY")
if not access_key or len(sys.argv) not in (2, 3):
raise SystemExit(
"Set SCREENSHOTONE_ACCESS_KEY, then run: "
"python3 url_to_pdf.py URL [output.pdf]"
)
target = sys.argv[1]
output = Path(sys.argv[2] if len(sys.argv) == 3 else "website.pdf")
payload = {
"url": target,
"format": "pdf",
"pdf_paper_format": "a4",
"pdf_margin": "12mm",
"pdf_print_background": True,
"media_type": "print",
}
request = Request(
"https://api.screenshotone.com/take",
data=json.dumps(payload).encode("utf-8"),
headers={
"Content-Type": "application/json",
"X-Access-Key": access_key,
},
method="POST",
)
try:
with urlopen(request, timeout=120) as response:
pdf = response.read()
except HTTPError as error:
raise SystemExit(f"PDF request failed: HTTP {error.code}") from None
except (URLError, TimeoutError) as error:
raise SystemExit(f"PDF request didn't complete: {error}") from None
if not pdf.startswith(b"%PDF-"):
raise SystemExit("The response isn't a PDF; nothing was written.")
output.write_bytes(pdf)
print(f"Saved {output} ({len(pdf)} bytes)")
if __name__ == "__main__":
main()

Run it:

Terminal window
python3 url_to_pdf.py 'https://example.com/' example.pdf

Both scripts hold the whole PDF in memory before writing it. That’s fine for normal pages. If you’re archiving thousands of large documents, have the API upload them straight to S3-compatible storage instead.

Choose the PDF layout

These options make the biggest difference to the result:

You wantSet
A printable A4 reportpdf_paper_format: "a4", media_type: "print", a margin
US Letterpdf_paper_format: "letter"
Landscapepdf_landscape: true
Background colors and imagespdf_print_background: true
The page as it looks on screenmedia_type: "screen"
Everything on one long pagepdf_fit_one_page: true, usually with screen styles

For a paginated document, set the paper format explicitly instead of relying on the default. For one long page with pdf_fit_one_page: true, remove pdf_paper_format from the examples and set pdf_margin: "0" so a fixed paper size or margins don’t interfere with the calculated page dimensions. The full list is in the PDF options reference.

The media_type choice surprises people the most. Many sites ship print CSS that hides the navigation, swaps colors, and adds page breaks. That often gives you a better document, even though it doesn’t look like the browser tab. If you want a visual record of the page as people see it, use screen.

Also, full_page=true is a screenshot option and doesn’t apply to PDFs. Use the pdf_* options above.

Wait for the content

If your PDF shows an empty dashboard or a spinner, the capture happened before the data loaded. Tell the API what “ready” looks like:

{
"wait_for_selector": "[data-report-ready]",
"delay": 1
}

A fixed delay on its own is a guess. A selector that only appears once the data is in works much better when load times vary. Charts, fonts, and images all finish at different moments, so open the exported file and check it. Don’t assume the page’s load event means the report is done.

If you control the page, a bit of print CSS goes a long way. Use break-inside: avoid on cards and table rows so they don’t split across pages, and add explicit breaks between sections. Just don’t put break-inside: avoid on anything taller than a page, because it can’t fit anywhere.

When the PDF looks wrong

ProblemWhat to check
A login screenThe renderer isn’t signed in. Pass a session explicitly; see the request options
Blank or half-drawn chartsWait for a selector that marks the page as ready
No background colorsTurn on pdf_print_background
Odd navigation or page breaksCompare media_type=print with media_type=screen
Tiny text on one huge pageDrop pdf_fit_one_page and let it paginate
TimeoutsCheck what the page is loading, or switch to async rendering

If you add retries, only retry failures that can recover, such as temporary network or rendering errors. Retrying a bad key or an invalid option won’t fix anything.

Read more Screenshot rendering

Interviews, tips, guides, industry best practices, and news.

View all posts

Automate website screenshots

Exhaustive documentation, ready SDKs, no-code tools, and other automation to help you render website screenshots and outsource all the boring work related to that to us.