Last updated: 2026-06-29 (end of rendering + rasterize-perf + interruptible-render session)
- Branch:
main@1c6e089, pushed toorigin/main(in sync). - GitHub Pages deployed & verified live: https://lighttransport.github.io/nanopdf/
(deploy workflow
.github/workflows/deploy-wasm-pages.ymlsucceeded; uses the committedexamples/wasm/src/nanopdf.{js,wasm}+ a Vite JS build). - Feature branch
fix-rendering-fonts-images-shadingsalso on origin (== main tip). - Build dirs:
build/(Release, libnanopdf + tests),build-rasterize/(linksbuild/libnanopdf.avia the/home/syoyo/work/nanopdfsymlink — rebuild BOTH after a src change or you hit a stale binary),build_wasm/(emscripten;source /mnt/nvme02/work/emsdk/emsdk_env.shthencmake --build build_wasm). - All tests green:
cd build && ctest→ unit / integration / validation / visual_lightvg. - Primary test doc:
data/dcsdd.pdf(ACM/TOG paper, embedded Type1 fonts, shadings, soft masks, transparency groups, JPEG figures). Reference renders compared vspdftoppm.
- Correctness: fig-14 luminosity soft-mask gradation; dense shading-function sampling (gradient ramp shape); transparency-group constant-alpha ("ghost" panels fig 2/7); tx-fonts big delimiters (eq 4/5/8/9); Type1 RD-binary parse; soft-mask snapshot clip (also a perf win).
- Rasterize perf @3000px: worst page 2.62→0.62s, full doc
--all11.5→6.35s, all pixel-identical. Levers: soft-mask snapshot→clipper bbox (2.7–4× on soft-mask pages); parallel image pre-decode + parallel area-average resample; tiled parallel scene rasterization (mask-free pages, main canvas only); opt-in--fast-png(fpnge). - Interruptible render:
RenderProgressCallbackreturns bool (false=cancel, fires everyprogress_percent_step);RenderResult.interrupted; WASMnanopdf_render_page_budget(...,budget_ms); viewer budgets cold/adjacent previews.
See memory notes: project_rasterize_png_bottleneck, project_interruptible_render,
project_shading_sh_ctm_and_function, project_transparency_group_opacity,
project_tex_math_glyph_names, project_type1_interpreter_bugs.
Wall-clock @3000px is ~60-70% parse (parse_pdf_content), not the fill. Parse =
sequential content interpretation + per-Do draw_image (resample) + per-glyph Type1
charstring→outline→transformed-Shape building + render_soft_mask_group. Levers:
- Parallelize across images on image-heavy pages (page 8 = 1.27s, 34 JPEGs).
draw_imageruns sequentially perDo; the resample is parallel only within one image (parallel_for_rows), so 34 small images serialize. Defer image draws, resample+convert all in parallel, then composite in z-order. Risk: z-order / interleaving with vector content. Ceiling ~1.3-1.5× on page 8. Medium effort/risk. render_soft_mask_groupis 10-14% on soft-mask pages. It renders each of ~17 groups to a fresh offscreen (cappedkMaxSoftMaskGroupPixels=2 MP). Reuse a scratch buffer (member) to cut alloc/zero/page-fault churn; and/or lower the cap (masks are smooth gradients — near-identical OK) — but VERIFY fig-14 fade quality, it's the thing we just fixed. Low effort, low-moderate value.draw_imagearea-average resample: fixed-point reciprocal instead of 4 int divides/dst-pixel (±1 LSB; user OK'd near-identical). Small (~5%). Low effort.
- Glyph bitmap cache for Type1 (rasterize once, blit) would cut the 16%
fill_polygon_aaon text pages, BUT the codebase deliberately bypasses bitmap caching for TJ body text (in_tj_text_draw_gate) to keep subpixel positioning. Enabling it snaps glyphs to the pixel grid — a visible text change. Only do with explicit sign-off; verify text crops vs current. High value, real quality risk.
- Raw-CFF simple fonts (
FontFile3) fail to load → substitute (math italic CM, Courier mono weight).cff_wrapper::wrap_cff_in_opentypeemits OTTO with only theCFFtable; stbtt/ttf_parse need head/hhea/maxp/hmtx(+cmap). Synthesize those + select glyphs by name via the CFF charset. Seeproject_tex_math_glyph_names→ "Biggest remaining opportunity". Larger cross-cutting change; affects other PDFs.
--fast-png(fpnge) is opt-in (files ~1.9× larger, lossless). If batch throughput matters more than size, consider a size-budgeted auto mode. Default stays miniz (smaller). Done as opt-in; revisit only on request.
- Cancel is checked only at
percent_stepboundaries during parse; NOT mid- image-decode, mid-render_soft_mask_group, or during the parallel pre-decode barrier. For a single very heavy operator the budget can overshoot. Add checkpoints insidedraw_image/render_soft_mask_groupif needed. Low value unless a page has one dominant op. - WASM cancel is time-budget only (single-threaded JS can't set a flag mid-render).
For true scroll-driven cancel, move rendering to a Web Worker + SharedArrayBuffer
cancel flag (needs COOP/COEP headers on Pages) and have the budget callback also
read the flag.
nanopdf_last_render_interrupted()already exists. Larger; only if the time-budget UX proves insufficient. --fast-pngis not wired into the WASM viewer (CLI-only). Not needed for the viewer (it paints pixels, doesn't PNG-encode), but exposing it for the viewer's "Export/Save as PNG" paths could speed those. Low.
- Cross-page parallelism (render N pages on N threads): ~3-4× for
--all, but the globalRenderCacheLRU singleton causes cross-page output divergence on a few image/shading pages (documented gotcha) and needs per-threadPdf. User previously chose single-page latency, so this was skipped. Revisit only if batch throughput becomes the goal; requires per-thread caches for determinism.
- Everything was tuned/verified against
data/dcsdd.pdfonly. Regression-render a wider corpus (otherdata/*.pdf, the math-font PDFs) to catch fonts/shadings the session's changes might affect — especially the dense shading-stop sampling and the transparency-groupgroup_opacity. Cheap, worth doing before further changes.
# build (rebuild BOTH or hit the stale-rasterize trap)
cmake --build build --target nanopdf -j8
cmake --build build-rasterize --target rasterize -j8
# perf baseline @3000px (TGA = render only; PNG includes encode)
for p in 1 7 8 5; do ./build-rasterize/rasterize data/dcsdd.pdf /tmp/p$p.tga -p $p -w 3000 --backend lightvg --format tga; done
# all pages -> compare_output/dcsdd_p300 (300 dpi)
for p in $(seq 1 14); do ./build-rasterize/rasterize data/dcsdd.pdf compare_output/dcsdd_p300/page_$(printf %02d $p).png -p $p --dpi 300; done
# pixel-diff a change vs prior binary: git stash, rebuild, render base, pop, rebuild, render new, compare in python
# tests
cd build && ctest --output-on-failure
# WASM (after src change): touch the .cc, rebuild, it auto-copies to examples/wasm/src/
source /mnt/nvme02/work/emsdk/emsdk_env.sh && cmake --build build_wasm -j8
# live viewer: vite dev on :5173 (examples/wasm), or the deployed Pages sitePerf method that worked: perf record -g --call-graph dwarf then perf report --no-children; but for WALL-CLOCK splits use direct std::chrono instrumentation —
perf sums parallel-thread CPU and over-counts the parallel decode/resample.
Continue optimizing/maintaining nanopdf rendering. State:
main@1c6e089is pushed and live on GitHub Pages. This session fixed dcsdd.pdf rendering (soft-mask gradation, transparency-group opacity, tx-fonts delimiters, Type1) and made rasterize ~2× faster at 3000px (soft-mask snapshot clip, parallel image decode/resample, tiled rasterization, opt-in fpnge) plus added interruptible rendering (bool progress callback + WASM time-budget render + viewer preview budgets). All pixel-identical, all tests green. Seetask.md"Remaining tasks" and theproject_*memory notes. Next, I want to : (A) parallelize per-image draw on image-heavy pages (page 8); (B) reuse/shrink soft-mask group offscreen buffers; (C) load raw-CFF FontFile3 simple fonts (math italic / mono weight) instead of substituting; (D) regression-render a wider PDF corpus to validate this session's changes; (E) move WASM rendering to a Web Worker + SharedArrayBuffer for true scroll-cancel. Rebuild BOTH build/ and build-rasterize/ (stale-binary trap), verify pixel-identity vs the prior binary, and run ctest.