The VPS tax on web scraping

If you've ever built a scraper or an automated testing pipeline with Puppeteer or Playwright, you know the drill. You spin up a VPS — usually a $5–$20/mo Hetzner, DigitalOcean, or Linode box — install Node, install Chrome dependencies, fight with libX11 and libnss3, and pray your headless browser doesn't eat all the RAM and crash at 3am.

I've done this across multiple projects. The pattern is always the same:

  1. Provision a VPS with enough RAM for Chrome (minimum 1GB, realistically 2GB+)
  2. Install system dependencies (apt-get install a dozen packages)
  3. Set up a Node process with Puppeteer or Playwright
  4. Add a process manager (PM2, systemd) so it restarts on crash
  5. Monitor memory usage because Chrome leaks like a sieve
  6. Handle zombie processes when browser.close() fails silently
  7. Pay monthly whether you scrape 10 pages or 10,000
The real cost. It's not just the $10/mo server bill. It's the ops overhead: updating Chrome, patching the OS, restarting crashed processes, debugging memory leaks at 2am. You wanted to scrape data — instead you're a sysadmin.

For my coupon scraping projects I was running Puppeteer on a Hetzner VPS. It worked, but I was babysitting it constantly. Chrome would occasionally hang, the process would eat 1.5GB of RAM on a 2GB box, and I had cron jobs to kill -9 zombie Chrome processes every hour. There had to be a better way.

Enter Cloudflare Browser Rendering

Cloudflare has a service called Browser Rendering that gives you a full headless Chromium instance inside a Worker. No VPS. No Docker. No Chrome installation. You write a Worker, call the Puppeteer API, and Cloudflare spins up a browser session on their infrastructure.

The architecture is dead simple: Trigger (Cron / HTTP) → Worker (your code) → Browser (Chromium) → Target (any website). Each request gets its own isolated browser session, spun up on demand and torn down when you're done.

You get full Puppeteer API access — page.goto(), page.evaluate(), page.screenshot(), page.waitForSelector() — all running on Cloudflare's edge network. And since it's a Worker, you get cron triggers, D1 for storage, KV for caching, and R2 for screenshots. The entire scraping pipeline lives in one place.

Key insight. Browser Rendering isn't a Puppeteer wrapper or a proxy. It's an actual Chromium instance running in Cloudflare's infrastructure. You use the real @cloudflare/puppeteer package, which is a fork of Puppeteer that connects to their browser pool instead of a local Chrome installation.

A real example: scraping coupon codes

Here's a stripped-down version of what a scraping Worker looks like. This fetches coupon codes from a store page, extracts structured data, and stores it in D1:

import puppeteer from "@cloudflare/puppeteer";

export default {
  async scheduled(controller, env, ctx) {
    // Launch browser from Cloudflare's pool
    const browser = await puppeteer.launch(env.BROWSER);
    const page = await browser.newPage();

    // Navigate and wait for content to load
    await page.goto("https://example.com/store/nike", {
      waitUntil: "networkidle0",
    });

    // Extract coupons using page.evaluate
    const coupons = await page.evaluate(() => {
      return [...document.querySelectorAll(".coupon-card")].map(el => ({
        code: el.querySelector(".code")?.textContent,
        discount: el.querySelector(".discount")?.textContent,
        expiry: el.querySelector(".expiry")?.textContent,
      }));
    });

    // Store in D1
    for (const coupon of coupons) {
      await env.DB.prepare(
        `INSERT OR REPLACE INTO coupons (code, discount, expiry)
         VALUES (?, ?, ?)`
      ).bind(coupon.code, coupon.discount, coupon.expiry).run();
    }

    await browser.close();
  },
};

And the wrangler.toml config:

# wrangler.toml
name = "coupon-scraper"
main = "src/index.ts"
compatibility_date = "2025-01-01"

# Browser Rendering binding
[browser]
binding = "BROWSER"

# D1 database for storage
[[d1_databases]]
binding = "DB"
database_name = "coupon-db"
database_id = "your-d1-id"

# Run every 4 hours
[triggers]
crons = ["0 */4 * * *"]

That's the entire scraping infrastructure. No server to provision. No Chrome to install. No process manager. Deploy with npx wrangler deploy and it runs on schedule.

VPS vs. Cloudflare Workers: the full comparison

VPS + Puppeteer Cloudflare Workers
Setup time 30–60 minutes (OS, deps, Chrome) 5 minutes (wrangler init + deploy)
Chrome installation Manual (apt-get + dependencies) Managed by Cloudflare
Scaling Manual (bigger VPS or multiple servers) Automatic (runs at edge, 300+ cities)
Memory management Your problem (Chrome leaks are real) Cloudflare's problem
Process crashes You handle restarts (PM2, systemd) Automatic per-request isolation
Cost (light use) $5–20/mo fixed (even idle) Free tier: 30 browser sessions/day
Cost (heavy use) $20–50/mo for 2–4GB RAM $0.20 per 1,000 sessions (paid plan)
Cron scheduling System crontab + custom scripts Built-in cron triggers
Storage Separate DB setup (Postgres, SQLite) D1, KV, R2 all integrated
Deployment SSH + git pull + restart npx wrangler deploy
Zombie processes Frequent (needs kill scripts) Impossible (request-scoped)
OS updates Your responsibility Not your problem

Advanced patterns

1. Reuse browser sessions

Cold-starting a browser is expensive. Cloudflare lets you keep sessions alive and reconnect to them, which is perfect for scraping multiple pages in sequence:

// Keep browser session alive for reuse
const browser = await puppeteer.launch(env.BROWSER, {
  keep_alive: 60000, // Keep alive for 60 seconds
});

// Later, reconnect to the same session
const browser = await puppeteer.connect(env.BROWSER, sessionId);

2. Screenshots to R2

Need visual proof of what you scraped? Screenshots go straight to R2 object storage:

const screenshot = await page.screenshot({ type: "png" });
await env.SCREENSHOTS.put(
  `scrape-${Date.now()}.png`,
  screenshot
);

3. Combine with AI

This is where it gets interesting. Cloudflare Workers AI runs on the same infrastructure. You can scrape a page and feed it straight into an LLM for extraction:

// Scrape the page content
const html = await page.content();

// Feed it to Workers AI for structured extraction
const result = await env.AI.run("@cf/meta/llama-4-scout-17b-16e-instruct", {
  messages: [{
    role: "user",
    content: `Extract all coupon codes from this HTML as JSON:\n${html}`
  }]
});

No API keys to manage. No external LLM calls. The browser, the AI model, and the database all live in the same Worker. This is the real power move — a complete scrape-extract-store pipeline in a single deployment.

The gotchas (be honest)

It's not all sunshine. There are real limitations to know about:

When to still use a VPS

I'm not saying burn your servers. There are cases where a VPS still wins:

For everything else — scheduled scraping, data extraction, screenshot generation, testing, monitoring — Workers + Browser Rendering is the better tool.

The migration is simple

If you're already using Puppeteer, the migration is a near drop-in replacement. The API is intentionally compatible:

// Before (VPS)
import puppeteer from "puppeteer";
const browser = await puppeteer.launch();

// After (Cloudflare Worker)
import puppeteer from "@cloudflare/puppeteer";
const browser = await puppeteer.launch(env.BROWSER);

One import change. One argument change. That's the diff.

Your page.goto(), page.evaluate(), page.click(), page.type() — all the same. The mental model is identical. You just stopped being a sysadmin.

Bottom line. A VPS for Puppeteer is like renting an apartment to store a suitcase. You're paying for space, utilities, and maintenance for something that only runs a few minutes per day. Cloudflare Workers let you pay for those few minutes and nothing else — while someone else handles the plumbing.

I've moved all my scraping pipelines to Workers. The coupons system that used to crash twice a week on Hetzner now runs on a cron trigger with zero maintenance. If you're still SSH-ing into servers to restart Chrome, try this. You won't go back.

← All writing Home