To convert a website to Markdown programmatically, render the page, extract its content, and save the Markdown response to a file. With ScreenshotOne, you request format=markdown through the same API used for website rendering.
This guide covers curl, JavaScript, and Python, including pages whose content appears after JavaScript runs. If you only need to convert one URL in your browser, use the free URL-to-Markdown tool.
What website-to-Markdown conversion produces
Markdown makes page content easier to store in notes, process in a script, or pass to a search or LLM pipeline. Depending on the source document, the result can contain headings, paragraphs, links, lists, and image references.
It is a text representation of the rendered document. It does not preserve the site’s typography, responsive layout, charts, or interactive behavior. Images can be represented by links; their visual contents are not automatically transcribed into text.
ScreenshotOne removes several non-content elements during Markdown conversion, including scripts and styles. That helps with readability, but it is not a guarantee that every site’s navigation, footer, or duplicated text will disappear. Inspect a sample from your target site before relying on the output.
Conversion also does not summarize the page or verify its claims. If your application needs the visual context, preserve a screenshot alongside the Markdown.
Get an API access key
Create a ScreenshotOne account and copy its API access key. Set it in the environment used by your server-side script:
export SCREENSHOTONE_ACCESS_KEY="your_access_key"The examples authenticate with the X-Access-Key header. This keeps the key out of the target URL and the generated Markdown. Keep these requests on your server or in a local script, and avoid putting the key into a public frontend bundle.
Convert a URL to Markdown with curl
Start with a public page:
curl --fail-with-body --get "https://api.screenshotone.com/take" \ --header "X-Access-Key: $SCREENSHOTONE_ACCESS_KEY" \ --data-urlencode "url=https://example.com/" \ --data-urlencode "format=markdown" \ --max-time 120 \ --output page.mdOn success, page.md contains the Markdown response directly. There is no JSON wrapper to decode in this request. Check curl’s exit status: a failed request can leave an error body in the output file, which should not be treated as converted page content.
The format option controls the response output. The separate markdown request option supplies Markdown as input for rendering an image or PDF; use format=markdown when you want to extract Markdown from a URL.
Convert a webpage to Markdown with JavaScript
This example uses Node.js 22.19 or later and its built-in fetch. Save it as url-to-markdown.mjs:
import { writeFile } from "node:fs/promises";
const accessKey = process.env.SCREENSHOTONE_ACCESS_KEY;if (!accessKey) { throw new Error("Set SCREENSHOTONE_ACCESS_KEY before running this script");}
const requestUrl = new URL("https://api.screenshotone.com/take");requestUrl.search = new URLSearchParams({ url: "https://example.com/", format: "markdown",}).toString();
const response = await fetch(requestUrl, { headers: { "X-Access-Key": accessKey }, signal: AbortSignal.timeout(120_000),});if (!response.ok) { throw new Error(`Markdown request returned HTTP ${response.status}`);}
const markdown = await response.text();if (!markdown.trim()) { throw new Error("The page returned empty Markdown");}await writeFile("page.md", markdown, "utf8");console.log("Saved page.md");Run it with:
node url-to-markdown.mjsChange the target url for your own page. URLSearchParams encodes the URL safely, including its query string, so you do not have to build the API request with string concatenation.
Convert a webpage to Markdown with Python
Install requests:
python -m pip install requestsSave this as url-to-markdown.py:
import osfrom pathlib import Path
import requests
response = requests.get( "https://api.screenshotone.com/take", headers={"X-Access-Key": os.environ["SCREENSHOTONE_ACCESS_KEY"]}, params={ "url": "https://example.com/", "format": "markdown", }, timeout=120,)response.raise_for_status()markdown = response.content.decode("utf-8")if not markdown.strip(): raise RuntimeError("The page returned empty Markdown")
Path("page.md").write_text(markdown, encoding="utf-8")print("Saved page.md")Run it with python url-to-markdown.py. The script checks the response before writing the file so an HTTP error is not silently saved as page content.
Convert JavaScript-rendered pages to Markdown
A JavaScript application may create its content after the document has loaded. Rendering the URL gives the application a browser, but you should also identify when the content you want is ready.
For example, the Quotes to Scrape demo inserts quote cards with JavaScript. Request those rendered cards with:
curl --fail-with-body --get "https://api.screenshotone.com/take" \ --header "X-Access-Key: $SCREENSHOTONE_ACCESS_KEY" \ --data-urlencode "url=https://quotes.toscrape.com/js/" \ --data-urlencode "format=markdown" \ --data-urlencode "wait_for_selector=.quote .text" \ --max-time 120 \ --output quotes.mdFor the Python or JavaScript scripts above, replace their request parameters with the same url, format, and wait_for_selector values.
The selector wait checks for DOM presence. If your application inserts an empty container first, waiting for that container is too early. Choose an element or attribute that appears when its text is populated, such as .report[data-ready="true"] on a page you control.
The demo’s first quote is a useful readiness signal for its initial list. For progressively loaded content, wait for completion and define how many records you need. Navigation events and a fixed delay alone do not guarantee that an application has finished updating.
See the wait options for the available controls. If you need custom clicks, scrolling, or a local browser session before extraction, the Python JavaScript-page scraping tutorial shows how to work with the rendered DOM directly.
Get a screenshot and Markdown from one render
For page research, reports, or archiving, readable text and an image can be useful together. Use an image output format and request Markdown as content metadata:
curl --fail-with-body --get "https://api.screenshotone.com/take" \ --header "X-Access-Key: $SCREENSHOTONE_ACCESS_KEY" \ --data-urlencode "url=https://example.com/" \ --data-urlencode "format=png" \ --data-urlencode "response_type=json" \ --data-urlencode "metadata_content=true" \ --data-urlencode "metadata_content_format=markdown" \ --max-time 120 \ --output capture.jsonThe Markdown is now available through content.url in the JSON response. It is not the response body or a summary of the screenshot. Download it and the screenshot promptly if you need to keep them.
The screenshot and Markdown documentation includes complete Python and JavaScript examples that save both files and check for missing results.
Process multiple URLs deliberately
These examples convert one page. For a list of URLs, process each one and save its source URL alongside the resulting file. Use a stable filename or identifier so reruns do not overwrite a different page by accident.
Keep concurrency within your account’s limit, give every request a timeout, and record failures separately from successful conversions. Avoid unbounded automatic retries. A retry policy should account for the status code, your usage allowance, and whether the source page has changed.
When building a search or LLM pipeline, preserve the retrieval time and inspect the generated Markdown before chunking or indexing it. Conversion does not make external page content authoritative or trustworthy.
Common conversion problems
| Problem | Practical next step |
|---|---|
| Missing JavaScript content | Add a selector that represents ready content and verify it against the page. |
| Navigation or footer text remains | Inspect the result and apply cleanup appropriate to that site’s document structure. |
| A graph or image loses its meaning | Save a screenshot for visual context; text conversion cannot preserve every visual feature. |
| Only the first part of a list appears | Check pagination, lazy loading, and whether the application virtualizes the DOM. |
| A returned content link stops working | Download temporary metadata files before their expiration time. |
| You need specific structured fields | Request HTML and parse it with Cheerio or another HTML parser. |
For a single conversion without writing code, open the URL-to-Markdown tool. For automation, start with one of the scripts above and test it against a representative page before processing a larger collection.
Frequently Asked Questions
If you read the article, but still have questions. Please, check the most frequently asked. And if you still have questions, feel free reach out at support@screenshotone.com.
How do I convert a website to Markdown with Python?
Request the webpage through ScreenshotOne with format=markdown, check the HTTP response, and save its UTF-8 content to a .md file. For JavaScript-generated content, add a readiness condition such as wait_for_selector.
Can I convert a JavaScript-rendered webpage to Markdown?
Yes. Render the page before conversion and wait for the relevant content. A browser rendering API can execute the page's scripts, but you still need a readiness condition when the application loads content later.
Is converting a page to Markdown the same as summarizing it?
No. Conversion changes the representation of extracted page content. It does not write a summary, verify the page's claims, or preserve the exact visual layout.
Can I get Markdown and a screenshot from the same page?
Yes. Request an image format with metadata_content=true and metadata_content_format=markdown. The JSON response supplies a screenshot URL and a temporary content URL that you can download.


