Save Now, Scrape Later
With Gaffa's capture_snapshot action, you can save an offline copy of any web page and run parse_json, generate_markdown and more against it later, even after the original page has changed or disappeared.
Sep 24, 2026

You don't always know what data you'll need from a webpage until later. By then, the page might have changed, been taken down, or moved behind a login. And while the Wayback Machine can sometimes help, it won't always have the exact version of the page you saw, or the parts of it you actually care about. With Gaffa, you don't have to guess ahead of time what you'll extract. With capture_snapshot, you can save a static copy of a page the moment you see it, then decide later exactly what to pull out and in what format.
What capture_snapshot actually gives you
capture_snapshot returns a single HTML file that captures the page as it looked at the moment of the request, with images and CSS embedded. It opens offline, on any machine, with no connection to the original site required. The saved file disables JavaScript, so anything interactive is frozen exactly as it was when you captured it.
That last part matters for what comes next. Because the output is a normal HTML file with a URL of its own, you can feed that URL straight back into Gaffa as the target of a brand new request, and run any extraction action against it whenever you're ready.
Gaffa has a demo article page you can use to try this yourself, one that generates a fresh set of paragraphs and images on every visit, so it's a decent stand-in for a page you know is going to look different by the time you come back to it. Here's a request that saves a snapshot of it:
Capturing a Snapshot
Running that returns a URL pointing to the saved HTML file, something like this:
Output from Capturing a Snapshot
Save that URL; it's what makes everything below possible.
Also note: How long the URL stays accessible depends on your plan's data retention period, so if you plan to come back to a snapshot weeks or months later, check that your plan covers it before you rely on it.
Decide how to extract the data later
Imagine you're tracking product pages, articles, or listings you might want structured data from eventually, but you're not sure yet what that structure should look like. Instead of building a schema before you've seen the full picture, you can snapshot pages as you come across them and work out the extraction later, without needing the original page to still exist or look the same.
When you're ready, pass the snapshot's URL as the URL of a new request and run parse_json against it, the same way you would against a live page:
Parsing a Saved Snapshot Into JSON
parse_json runs on the gpt-4o-mini model by default, and it doesn't care whether the page you're pointing it at is live on the web or sitting in a snapshot file; it reads the content either way.
Change your extraction format without revisiting the page
The same snapshot URL isn't locked to one action either. If you parse a batch of pages into JSON and later realise you actually want the readable text instead, you don't need to go back to the original site to fix it. Feed the same URL into a new request with generate_markdown instead:
Converting a Saved Snapshot to Markdown
Because you're loading a lightweight static file rather than the live page, there's no navigation, ads, or third-party scripts for Gaffa to load first, so these requests tend to complete faster than the original capture did, and typically use fewer credits too, since there's far less for the browser to fetch and render.
What kinds of actions work on a snapshot URL
parse_json and generate_markdown aren't special cases here. Any action that just reads the page content already sitting in front of it works the same way against a snapshot. parse_table finds a table in the saved HTML exactly as it would on a live page, capture_screenshot renders the static file and takes the shot, print exports it to PDF the same way it would export any other page, and capture_dom and generate_simplified_dom both work too, since they're only reading and reshaping the DOM that's already there.
What doesn't work are actions that need to interact with a live, running page, things like scroll, click, or type. JavaScript is switched off in a saved snapshot, so nothing is left to trigger. If a page needs interaction, that has to happen before you capture it, not after.
Load everything before you capture
capture_snapshot saves the page exactly as it appears when you take the shot, which means anything that hasn't loaded yet simply won't be there. Gaffa's demo storefront can be configured to load its product listings as you scroll, using the loadingMode, itemCount, and itemLoadTime parameters in its URL. If you captured a snapshot without reaching the bottom of the page, those later products wouldn't be in the saved file, and no amount of waiting in place would load them, since they load only in response to scrolling.
Adding a wait for the first product to render, followed by a scroll action, takes care of this: the wait makes sure the grid has actually mounted before anything tries to move the page, and the scroll then takes you to the bottom and keeps going for as long as new items keep appearing, so you don't need to guess how many there'll be.
Scrolling to Load Every Product Before Capturing
The wait step waits for the first product card to appear, so the scroll action never runs against an empty, unrendered page. From there, percentage: 100 sends the browser to the bottom; timeout: 5000 gives Gaffa up to 5 seconds to treat the page as scrollable rather than giving up the instant it checks; and wait_time: 1000 does the lazy-loading work. After each arrival at the bottom, Gaffa watches for a second, and if the page has grown, it keeps scrolling and the wait resets. max_scroll_time caps the whole thing at 25 seconds, so a page that never stops loading (as this one can, with itemCount=infinite) doesn't run the action forever; it just stops cleanly rather than failing.
Since JavaScript doesn't run inside the saved file, this is also your only chance to expand collapsed sections, open tabs, or dismiss cookie banners using an action like click before you capture. Anything you don't open beforehand stays closed in the snapshot.
Other things you can do with a snapshot
Extraction is only one use for a saved page. A snapshot is also a straightforward way to keep proof of what a page looked like on a given date, whether that's for a pricing dispute, a compliance record, or just your own reference before a competitor's site inevitably redesigns itself. Since the file opens offline in any browser, you can hand it to someone else, store it alongside other records, or drop it into a shared drive without requiring a Gaffa account.
Capturing a page doesn't have to mean deciding on the spot what you'll do with it. With capture_snapshot, you can save the page now and work out the extraction, schema, or format later, without needing the original site to still be there. If you haven't already, sign up at gaffa.dev and try it in the API Playground, or see how to slash your Gaffa credit costs with other ways to get more out of every request.
Frequently asked questions
What does capture_snapshot actually save?
It saves a single offline HTML file of the page as it looked at capture time, with all images and CSS embedded. JavaScript is disabled, so anything interactive stays frozen as it was.
Can I run parse_json against a saved snapshot instead of the live page?
Yes. Pass the snapshot's URL as the URL in a new request and run parse_json against it exactly as you would a live page or online PDF, with your schema defined the same way.
Can I convert a saved snapshot to markdown instead of JSON?
Yes, use generate_markdown with the snapshot's URL as the request's URL. You can run either action, or both, against the same saved snapshot at any time.
What other actions can I run against a saved snapshot?
Any action that just reads the rendered page works, including parse_table, capture_screenshot, print, capture_dom, and generate_simplified_dom. Actions that need to interact with a live page, like scroll or click, won't do anything.
Why is content missing from a snapshot?
The page probably loads that content dynamically, after a delay or as the user scrolls. Add a wait or scroll before capture_snapshot so everything has rendered before the capture happens.
How long does a snapshot's URL stay accessible?
That depends on your plan's data retention period, for example, 7 days on the Starter plan, 30 days on Startup, or 3 months on Growth. Extract what you need from it before that window runs out.
Can I use capture_snapshot just to archive a page, without extracting anything?
Yes. The saved HTML file opens offline on any machine, which makes it useful for keeping a record of how a page looked on a given date, independent of any extraction step.
