The VPS tax on web scraping
If you've ever built a scraper or an automated testing pipeline with Puppeteer or Playwright, you know the drill. You spin up a VPS — usually a $5–$20/mo Hetzner, DigitalOcean, or Linode box — install Node, install Chrome dependencies, fight with libX11 and libnss3, and pray your headless browser doesn't eat all the RAM and crash at 3am.
I've done this across multiple projects. The pattern is always the same:
- Provision a VPS with enough RAM for Chrome (minimum 1GB, realistically 2GB+)
- Install system dependencies (
apt-get installa dozen packages) - Set up a Node process with Puppeteer or Playwright
- Add a process manager (PM2, systemd) so it restarts on crash
- Monitor memory usage because Chrome leaks like a sieve
- Handle zombie processes when
browser.close()fails silently - Pay monthly whether you scrape 10 pages or 10,000
The real cost. It's not just the $10/mo server bill. It's the ops overhead: updating Chrome, patching the OS, restarting crashed processes, debugging memory leaks at 2am. You wanted to scrape data — instead you're a sysadmin.
For my coupon scraping projects I was running Puppeteer on a Hetzner VPS. It worked, but I was babysitting it constantly. Chrome would occasionally hang, the process would eat 1.5GB of RAM on a 2GB box, and I had cron jobs to kill -9 zombie Chrome processes every hour. There had to be a better way.
Enter Cloudflare Browser Rendering
Cloudflare has a service called Browser Rendering that gives you a full headless Chromium instance inside a Worker. No VPS. No Docker. No Chrome installation. You write a Worker, call the Puppeteer API, and Cloudflare spins up a browser session on their infrastructure.
The architecture is dead simple: Trigger (Cron / HTTP) → Worker (your code) → Browser (Chromium) → Target (any website). Each request gets its own isolated browser session, spun up on demand and torn down when you're done.
You get full Puppeteer API access — page.goto(), page.evaluate(), page.screenshot(), page.waitForSelector() — all running on Cloudflare's edge network. And since it's a Worker, you get cron triggers, D1 for storage, KV for caching, and R2 for screenshots. The entire scraping pipeline lives in one place.
Key insight. Browser Rendering isn't a Puppeteer wrapper or a proxy. It's an actual Chromium instance running in Cloudflare's infrastructure. You use the real @cloudflare/puppeteer package, which is a fork of Puppeteer that connects to their browser pool instead of a local Chrome installation.
A real example: scraping coupon codes
Here's a stripped-down version of what a scraping Worker looks like. This fetches coupon codes from a store page, extracts structured data, and stores it in D1:
import puppeteer from "@cloudflare/puppeteer";
export default {
async scheduled(controller, env, ctx) {
// Launch browser from Cloudflare's pool
const browser = await puppeteer.launch(env.BROWSER);
const page = await browser.newPage();
// Navigate and wait for content to load
await page.goto("https://example.com/store/nike", {
waitUntil: "networkidle0",
});
// Extract coupons using page.evaluate
const coupons = await page.evaluate(() => {
return [...document.querySelectorAll(".coupon-card")].map(el => ({
code: el.querySelector(".code")?.textContent,
discount: el.querySelector(".discount")?.textContent,
expiry: el.querySelector(".expiry")?.textContent,
}));
});
// Store in D1
for (const coupon of coupons) {
await env.DB.prepare(
`INSERT OR REPLACE INTO coupons (code, discount, expiry)
VALUES (?, ?, ?)`
).bind(coupon.code, coupon.discount, coupon.expiry).run();
}
await browser.close();
},
};
And the wrangler.toml config:
# wrangler.toml
name = "coupon-scraper"
main = "src/index.ts"
compatibility_date = "2025-01-01"
# Browser Rendering binding
[browser]
binding = "BROWSER"
# D1 database for storage
[[d1_databases]]
binding = "DB"
database_name = "coupon-db"
database_id = "your-d1-id"
# Run every 4 hours
[triggers]
crons = ["0 */4 * * *"]
That's the entire scraping infrastructure. No server to provision. No Chrome to install. No process manager. Deploy with npx wrangler deploy and it runs on schedule.
VPS vs. Cloudflare Workers: the full comparison
| VPS + Puppeteer | Cloudflare Workers | |
|---|---|---|
| Setup time | 30–60 minutes (OS, deps, Chrome) | 5 minutes (wrangler init + deploy) |
| Chrome installation | Manual (apt-get + dependencies) | Managed by Cloudflare |
| Scaling | Manual (bigger VPS or multiple servers) | Automatic (runs at edge, 300+ cities) |
| Memory management | Your problem (Chrome leaks are real) | Cloudflare's problem |
| Process crashes | You handle restarts (PM2, systemd) | Automatic per-request isolation |
| Cost (light use) | $5–20/mo fixed (even idle) | Free tier: 30 browser sessions/day |
| Cost (heavy use) | $20–50/mo for 2–4GB RAM | $0.20 per 1,000 sessions (paid plan) |
| Cron scheduling | System crontab + custom scripts | Built-in cron triggers |
| Storage | Separate DB setup (Postgres, SQLite) | D1, KV, R2 all integrated |
| Deployment | SSH + git pull + restart | npx wrangler deploy |
| Zombie processes | Frequent (needs kill scripts) | Impossible (request-scoped) |
| OS updates | Your responsibility | Not your problem |
Advanced patterns
1. Reuse browser sessions
Cold-starting a browser is expensive. Cloudflare lets you keep sessions alive and reconnect to them, which is perfect for scraping multiple pages in sequence:
// Keep browser session alive for reuse
const browser = await puppeteer.launch(env.BROWSER, {
keep_alive: 60000, // Keep alive for 60 seconds
});
// Later, reconnect to the same session
const browser = await puppeteer.connect(env.BROWSER, sessionId);
2. Screenshots to R2
Need visual proof of what you scraped? Screenshots go straight to R2 object storage:
const screenshot = await page.screenshot({ type: "png" });
await env.SCREENSHOTS.put(
`scrape-${Date.now()}.png`,
screenshot
);
3. Combine with AI
This is where it gets interesting. Cloudflare Workers AI runs on the same infrastructure. You can scrape a page and feed it straight into an LLM for extraction:
// Scrape the page content
const html = await page.content();
// Feed it to Workers AI for structured extraction
const result = await env.AI.run("@cf/meta/llama-4-scout-17b-16e-instruct", {
messages: [{
role: "user",
content: `Extract all coupon codes from this HTML as JSON:\n${html}`
}]
});
No API keys to manage. No external LLM calls. The browser, the AI model, and the database all live in the same Worker. This is the real power move — a complete scrape-extract-store pipeline in a single deployment.
The gotchas (be honest)
It's not all sunshine. There are real limitations to know about:
- Session limits on free tier. You get 30 browser sessions per day on the free plan. That's enough for light scraping or testing, but production workloads need the paid plan ($5/mo for Workers Paid, then $0.20 per 1,000 sessions).
- Execution time limits. Workers have a 30-second CPU time limit (paid plan). If your scrape involves heavy JavaScript rendering or many page navigations, you might hit it. Solution: break work into multiple Worker invocations using Durable Objects or queues.
- No persistent browser state. Each session starts fresh — no cookies, no localStorage from previous runs. If you need login persistence, you'll need to handle authentication in each session or store cookies in KV and restore them.
- Puppeteer API subset.
@cloudflare/puppeteercovers the most common APIs but isn't 100% feature-complete with upstream Puppeteer. Check the docs if you need niche features. No Playwright support yet — Puppeteer only. - Debugging is harder. On a VPS, you can SSH in and run Puppeteer interactively. With Workers, you're debugging through logs and
wrangler tail. The feedback loop is slightly longer.
When to still use a VPS
I'm not saying burn your servers. There are cases where a VPS still wins:
- Long-running browser sessions — crawling hundreds of pages in a single session over minutes or hours
- Playwright-specific features — if you need Firefox/WebKit rendering or Playwright-specific APIs
- Custom Chrome extensions — if your scraper loads extensions into the browser
- Residential proxy integration — some proxy providers work better with self-hosted Chrome
- Full browser profiles — if you need persistent login sessions across runs without re-authenticating
For everything else — scheduled scraping, data extraction, screenshot generation, testing, monitoring — Workers + Browser Rendering is the better tool.
The migration is simple
If you're already using Puppeteer, the migration is a near drop-in replacement. The API is intentionally compatible:
// Before (VPS)
import puppeteer from "puppeteer";
const browser = await puppeteer.launch();
// After (Cloudflare Worker)
import puppeteer from "@cloudflare/puppeteer";
const browser = await puppeteer.launch(env.BROWSER);
One import change. One argument change. That's the diff.
Your page.goto(), page.evaluate(), page.click(), page.type() — all the same. The mental model is identical. You just stopped being a sysadmin.
Bottom line. A VPS for Puppeteer is like renting an apartment to store a suitcase. You're paying for space, utilities, and maintenance for something that only runs a few minutes per day. Cloudflare Workers let you pay for those few minutes and nothing else — while someone else handles the plumbing.
I've moved all my scraping pipelines to Workers. The coupons system that used to crash twice a week on Hetzner now runs on a cron trigger with zero maintenance. If you're still SSH-ing into servers to restart Chrome, try this. You won't go back.