Ask Anvil

Answers to questions about automating PDFs, e-signatures, Webforms, and other paperwork problems.
PDFs
Categories

How do I generate thousands of PDFs from HTML without launching a browser for each one?

Calling puppeteer.launch() inside your per-document loop starts a new Chromium process for every PDF. Opening a page in a browser that is already running is much lighter, so the fix is to launch once, render each document in its own short-lived page, and cap how many pages run at the same time:

import puppeteer from 'puppeteer';

const CONCURRENCY = 4;   // pages rendering at the same time
const BATCH_SIZE = 500;  // relaunch Chromium after this many PDFs

async function renderBatch(browser, jobs, toHtml, save) {
  const queue = [...jobs];
  const failures = [];

  async function worker() {
    let job;
    while ((job = queue.shift()) !== undefined) {
      const page = await browser.newPage();
      try {
        await page.setContent(toHtml(job), { waitUntil: 'load', timeout: 30_000 });
        const pdf = await page.pdf({ format: 'A4', printBackground: true });
        await save(job, pdf);
      } catch (err) {
        failures.push({ job, error: err.message });
      } finally {
        await page.close();
      }
    }
  }

  await Promise.all(Array.from({ length: CONCURRENCY }, () => worker()));
  return failures;
}

export async function renderAll(jobs, toHtml, save) {
  const failures = [];
  for (let i = 0; i < jobs.length; i += BATCH_SIZE) {
    const browser = await puppeteer.launch();
    try {
      const batch = jobs.slice(i, i + BATCH_SIZE);
      failures.push(...(await renderBatch(browser, batch, toHtml, save)));
    } finally {
      await browser.close();
    }
  }
  return failures;
}

Call it with your records, a function that turns one record into HTML, and a function that stores the result:

import { mkdir, writeFile } from 'node:fs/promises';
import { renderAll } from './render.mjs';

const invoices = [
  { id: 'INV-1001', html: '<h1>Invoice INV-1001</h1><p>Total: $120.00</p>' },
  { id: 'INV-1002', html: '<h1>Invoice INV-1002</h1><p>Total: $86.50</p>' },
];

await mkdir('./out', { recursive: true });
const failures = await renderAll(
  invoices,
  (inv) => inv.html,
  (inv, pdf) => writeFile(`./out/${inv.id}.pdf`, pdf),
);
console.log(`done, ${failures.length} failed`, failures);

How it works

  • One browser per batch. Each job gets its own page, and the page is closed in a finally block so a template that throws never leaks an open tab.
  • CONCURRENCY workers pull from a shared queue. Because JavaScript runs one callback at a time, queue.shift() never hands the same job to two workers.
  • Failures are collected and returned instead of thrown, so one bad record does not stop a run of ten thousand. Retry the returned list afterwards.
  • BATCH_SIZE closes and relaunches Chromium every few hundred documents, so no single browser process runs long enough for gradual memory growth to pile up. Both numbers are starting points; tune them by measuring on your own hardware.
  • page.pdf() waits for document.fonts.ready by default and returns a Uint8Array, which writeFile accepts directly. The paper format defaults to Letter, so set it explicitly if you need A4.

Caveats

In current Puppeteer releases, setContent() only accepts 'load' or 'domcontentloaded' for waitUntil. Older versions (Puppeteer 21, for example) accepted 'networkidle0' here, but the current types reject it. 'load' waits for the stylesheets, scripts, and images referenced in the HTML (except lazy-loaded ones), but not for data a script fetches afterwards. If your template renders client-side, render it on the server first, or call page.waitForSelector() on an element that only appears once the data is in place before calling page.pdf().

Every open page holds its own DOM, layout, and images in memory, so a higher CONCURRENCY is not automatically faster. Start near the machine's CPU core count and watch memory as you raise it.

If you would rather not run Chromium on your own servers, a hosted HTML-to-PDF API such as Anvil's PDF generation API takes the same HTML and returns the PDF bytes, and the same queue, concurrency cap, and retry list apply to those calls.

Back to All Questions

The fastest way to build software for documents

Anvil Document SDK is a comprehensive toolbox for product teams launching document flows where PDF filling, signing, and complex conditional scenarios are necessary.
Explore Anvil
Anvil Webforms
Marketing Mode