How I Generate OG Images in 35ms Without a Headless Browser
Every dynamic OG image generator I could find — Vercel OG, Bannerbear, placid, the countless "og:image as a service" startups — runs a headless browser under the hood. That decision costs you ~500MB of RAM per worker, 500ms+ cold starts, and a dependency tree that breaks every six months.
I built OG Forge on a different bet: if you constrain the design space to templates instead of arbitrary HTML, you don't need a browser at all. A pure Pillow pipeline renders the same cards in ~35ms average, in a ~60MB Docker image that runs comfortably on any free tier.
Here's how the rendering works, how I profiled it down from 391ms, and the unit economics of running it as a freemium API.
The problem with browser-based OG renderers
When you curl an OG image endpoint, you're usually paying for:
- A Chromium instance (or a pool of them) staying warm
- Font loading, layout, paint, compositing — the full rendering pipeline, to produce what is ultimately a static image with a title, a subtitle, and maybe a logo
- Memory pressure that caps concurrency on small instances
For a card with two lines of text and a gradient, that's absurd overkill. The insight: template-based rendering (think "pick one of 5 layouts, fill in text, choose colors") covers the vast majority of real OG image needs. Docs sites, blogs, and changelogs don't need arbitrary HTML — they need consistently branded cards.
The rendering pipeline
OG Forge exposes one endpoint:
GET /v1/generate?title=Ship%20your%20side%20project&template=gradient&accent_color=%236366f1
A PNG comes out. Five templates — gradient, split, spotlight, banner, minimal — each with auto-fitting typography and full color control. Under the hood:
1. Auto-fitting typography
The naive approach (font = 64px; if too long, 48px) produces broken layouts. Instead, fit_font() binary-searches the largest size where the title still wraps into at most N lines, then wrap() reflows. Titles shrink and wrap — they never overflow or clip:
def fit_font(text, box_w, box_h, max_lines, path, start, end):
while start < end - 1:
mid = (start + end) // 2
font = ImageFont.truetype(path, mid)
lines = wrap(text, font, box_w)
if len(lines) <= max_lines and all(fits(l, font, box_w) for l in lines):
start = mid
else:
end = mid
return ImageFont.truetype(path, start), wrap(text, ..., box_w)
2. Gradients without full-size rotation
A diagonal gradient at 1200×630 needs a rotated linear gradient. Rotating a full-size image is expensive; instead, render the gradient on a small image, rotate that, then upscale with bilinear filtering. Identical visual result, a fraction of the pixel work.
3. Glows at quarter resolution
The spotlight template's soft radial glow is a blur — and blur cost scales with pixels. Rendering the glow at 1/4 resolution and upscaling is visually indistinguishable and 16x cheaper.
4. Gradient mask caching
Template masks are keyed by (size, angle) and cached in memory. Repeat renders skip the geometry work entirely.
The profiling story: the bottleneck wasn't where I expected
Naive benchmarks said "spotlight: 391ms". The obvious assumption is that rendering (the blur, the gradient math) dominates. Profiling said otherwise:
| Optimization | Before | After |
|---|---|---|
| Gradient: small-rotate + bilinear upscale | 391ms | ~180ms |
| Glow at 1/4 resolution | ~180ms | ~90ms |
| Mask caching | ~90ms | ~60ms |
PNG compress_level=6 instead of optimize=True | ~60ms | ~35ms |
The single biggest win was the last one — and it wasn't in rendering at all. It was in PNG encoding. Pillow's optimize=True sets compress_level=9, which spends milliseconds hunting for marginal byte savings. Dropping to compress_level=6 made encoding 4.4x faster at the cost of ~4KB per image. For an API that serves images over HTTP with gzip-in-transit, that's a straight win:
buf = io.BytesIO()
img.save(buf, "PNG", compress_level=6) # not optimize=True
Final numbers, 1200×630 output:
| Template | Render time |
|---|---|
| minimal | 18.3ms |
| banner | 21.8ms |
| split | 35.1ms |
| gradient | 42.3ms |
| spotlight | 58.5ms |
| average | 35.2ms |
The lesson I'd keep: profile before optimizing. I'd have spent a week micro-optimizing blur math that was 15% of the runtime, while a one-line PNG parameter was 40% of it.
The service architecture
It's a deliberately boring stack:
- FastAPI — async framework, auto docs at /docs (free developer experience win)
- Per-IP token bucket — a sliding window (deque of timestamps per hashed IP), 10 requests/min for the free tier, 429 with
Retry-Afterwhen exceeded - API-key tier hook —
X-API-Keyheaders resolve to a tier dict; in production a Creem webhook provisions keys automatically on payment (checkout.completed → grant, subscription.canceled → revoke). The rate limiter skips keyed requests - Edge caching — responses carry
Cache-Control: public, max-age=86400, so repeat renders of the same URL hit CDNs, not the origin
The whole thing is stateless. No database, no Redis (the rate-limit deque is per-worker and that's fine for a free tier), no workers. One container.
The unit economics of a freemium image API
Free tier marginal cost approaches zero: at 35ms per render, a single free-tier instance can produce ~28 images/second of headroom, and CDN caching absorbs repeat traffic. The paid tier ($9/mo for unlimited renders + priority + custom fonts) prices against the alternative cost — a $20-60/mo Bannerbear or a self-run Puppeteer fleet — not against marginal compute.
The self-host escape hatch matters commercially: the repo is MIT-licensed, so the free tier competes with... itself, self-hosted. That's intentional. Every self-hoster is a distribution channel (their og:image URLs point at their domain running your Docker image), and the conversion path is "got big enough to want managed, priority, custom fonts."
Full disclosure: an AI agent built this
The uncomfortable part of this post: I didn't write most of this code. OG Forge was built autonomously by an AI agent (GLM-5.3) in a single session — architecture, the five templates, the typography auto-fitting, and the entire profiling-driven optimization pass described above (it's the one that found the PNG encode bottleneck, not me). My role was direction, review, and shipping. The git history is public if you want to see the shape of it.
My honest take: "AI built it" is not the selling point of the product — 35ms renders are. But it's now true that the marginal cost of "idea → deployed, tuned, documented microservice" is a prompt away. The scarce inputs are judgment about what to build and the willingness to verify what comes out.
Try it
The playground needs no signup — the landing page calls the API live:
Playground: og-forge.xyz · API docs: /docs
curl "https://og-forge.xyz/v1/generate" \
--get --data-urlencode "title=Hello world" \
--data-urlencode "template=split" \
-o card.png
What template should be built next? I'm taking requests.