One of the most underestimated engineering problems in enterprise software is PDF generation. Every insurance claim produces an approval document. Every work order gets a client-facing summary. Every HRM system generates monthly payslips. At Edelta Corporation, this was happening manually — someone would download data, paste it into a Word template, and export to PDF.
I automated this entire process with a serverless PDF generation service on AWS Lambda. Here's the architecture.
The Requirements
Across the ProtectAll Claims system and the HRMs platform, we needed PDFs for:
- Claim approval letters (with claimant details, approved amount, policy number, assessor signature block)
- Inspection reports (with GPS location, photo attachments, field notes, inspector name)
- Work order summaries (itemized costs, contractor details, completion status)
- Payslips (employee name, department, gross/net pay, deductions breakdown)
Each had a different template but a common pattern: receive JSON data → render into a branded document → return a pre-signed S3 URL.
Architecture
Why Puppeteer + Handlebars?
Handlebars for templating — it's logic-minimal (good for templates), supports partials (shared headers/footers across document types), and the output is standard HTML/CSS which designers can style without knowing Node.js.
Puppeteer (headless Chromium) for HTML → PDF conversion — it produces the most accurate PDF output of any Node.js library because it's an actual browser rendering engine. Tables, CSS flexbox layouts, page breaks — all work exactly like a browser would display them.
The alternative (PDFKit, pdfmake) requires building documents programmatically, which is painful for designer collaboration and complex layouts.
Lambda Configuration Challenges
Puppeteer and Chromium add ~300MB to your Lambda package. The default Lambda limit is 250MB for zip deployment, but AWS Lambda container images support up to 10GB.
We deployed the PDF Lambda as a container image:
FROM public.ecr.aws/lambda/nodejs:20
# Install Chromium dependencies
RUN yum install -y \
at-spi2-atk \
atk \
cups-libs \
gtk3 \
libXcomposite \
nss \
...
# Install Puppeteer with bundled Chromium
ENV PUPPETEER_SKIP_DOWNLOAD=false
ENV PUPPETEER_CACHE_DIR=/tmp/.cache/puppeteer
COPY package*.json ./
RUN npm ci --omit=dev
COPY . .
CMD ["handler.handler"]
Using @sparticuz/chromium instead of the full Puppeteer Chromium gave us a Lambda-optimized build at ~45MB — back within zip deployment limits for most configurations.
The Template System
Templates live in an S3 bucket, not in the Lambda code. This means the design team can update templates without a code deployment:
// Fetch template from S3 at runtime
const templateSource = await s3.getObject({
Bucket: 'templates-bucket',
Key: `pdf-templates/${documentType}.hbs`,
}).promise();
const template = Handlebars.compile(templateSource.Body.toString());
const html = template(documentData);
Each template is a full HTML document with inline CSS (important — external CSS doesn't load reliably in headless Chromium when running in Lambda):
<!DOCTYPE html>
<html>
<head>
<style>
body { font-family: 'Arial', sans-serif; margin: 0; padding: 40px; }
.header { display: flex; justify-content: space-between; align-items: center; }
.claim-amount { font-size: 32px; font-weight: bold; color: #1a365d; }
@page { size: A4; margin: 20mm; }
</style>
</head>
<body>
<div class="header">
<img src="data:image/png;base64,{{logoBase64}}" alt="Logo" height="60" />
<div>
<h1>Claim Approval Letter</h1>
<p>Ref: {{claimReference}}</p>
</div>
</div>
<p>Dear {{claimantName}},</p>
<p>Your claim for <strong class="claim-amount">₹{{approvedAmount}}</strong>
has been approved on {{approvalDate}}.</p>
{{> signatureBlock assessorName=assessorName designation=designation}}
</body>
</html>
Images (logos, signatures) are base64-encoded and inlined — no external HTTP requests from inside Lambda.
Pre-Signed URL Strategy
After generating and uploading the PDF, we return a pre-signed S3 URL with a 24-hour expiry:
const signedUrl = await s3.getSignedUrlPromise('getObject', {
Bucket: process.env.PDF_BUCKET,
Key: s3Key,
Expires: 86400, // 24 hours
});
The caller (Zoho webhook handler) stores this URL in the CRM record. Claim officers and claimants can download the document via the URL without the document being publicly exposed on S3.
For long-lived documents (archived claims), a separate Lambda runs nightly to generate permanent URLs via CloudFront signed cookies.
Cold Start Mitigation
Puppeteer initialization is the expensive part (~2–3 seconds cold start). We kept the browser instance alive across invocations using module-level initialization:
// Initialize once per Lambda execution environment
let browser;
const getBrowser = async () => {
if (!browser || !browser.connected) {
browser = await puppeteer.launch({
executablePath: await chromium.executablePath(),
headless: chromium.headless,
args: chromium.args,
});
}
return browser;
};
This reuses the Chromium process for subsequent invocations within the same Lambda container, dropping warm invocation time from ~2.5s to ~200–400ms.
Metrics After 6 Months in Production
| Metric | Result |
|---|---|
| Documents generated per month | ~8,000 |
| Average generation time (warm) | 380ms |
| Average generation time (cold) | 2.8s |
| Lambda cost (estimated) | ~$3.20/month |
| Manual document prep time replaced | ~40 hours/month |
The cost efficiency of Lambda for this workload — bursty, relatively infrequent, latency-tolerant — is excellent. A dedicated EC2 instance for the same job would cost 20–30× more.
Automated PDF generation is one of those investments that pays back every week indefinitely. If your team is still manually formatting documents, this is the architecture to adopt.