Skip to content

Latest commit

 

History

History
147 lines (125 loc) · 8.95 KB

File metadata and controls

147 lines (125 loc) · 8.95 KB

nanopdf — task / handoff notes

Last updated: 2026-06-29 (end of rendering + rasterize-perf + interruptible-render session)

Current state

  • Branch: main @ 1c6e089, pushed to origin/main (in sync).
  • GitHub Pages deployed & verified live: https://lighttransport.github.io/nanopdf/ (deploy workflow .github/workflows/deploy-wasm-pages.yml succeeded; uses the committed examples/wasm/src/nanopdf.{js,wasm} + a Vite JS build).
  • Feature branch fix-rendering-fonts-images-shadings also on origin (== main tip).
  • Build dirs: build/ (Release, libnanopdf + tests), build-rasterize/ (links build/libnanopdf.a via the /home/syoyo/work/nanopdf symlink — rebuild BOTH after a src change or you hit a stale binary), build_wasm/ (emscripten; source /mnt/nvme02/work/emsdk/emsdk_env.sh then cmake --build build_wasm).
  • All tests green: cd build && ctest → unit / integration / validation / visual_lightvg.
  • Primary test doc: data/dcsdd.pdf (ACM/TOG paper, embedded Type1 fonts, shadings, soft masks, transparency groups, JPEG figures). Reference renders compared vs pdftoppm.

Done this session (11 commits, a352552..main)

  • Correctness: fig-14 luminosity soft-mask gradation; dense shading-function sampling (gradient ramp shape); transparency-group constant-alpha ("ghost" panels fig 2/7); tx-fonts big delimiters (eq 4/5/8/9); Type1 RD-binary parse; soft-mask snapshot clip (also a perf win).
  • Rasterize perf @3000px: worst page 2.62→0.62s, full doc --all 11.5→6.35s, all pixel-identical. Levers: soft-mask snapshot→clipper bbox (2.7–4× on soft-mask pages); parallel image pre-decode + parallel area-average resample; tiled parallel scene rasterization (mask-free pages, main canvas only); opt-in --fast-png (fpnge).
  • Interruptible render: RenderProgressCallback returns bool (false=cancel, fires every progress_percent_step); RenderResult.interrupted; WASM nanopdf_render_page_budget(...,budget_ms); viewer budgets cold/adjacent previews.

See memory notes: project_rasterize_png_bottleneck, project_interruptible_render, project_shading_sh_ctm_and_function, project_transparency_group_opacity, project_tex_math_glyph_names, project_type1_interpreter_bugs.


Remaining tasks / optimization opportunities (prioritized)

A. Rasterize performance (parse phase now dominates)

Wall-clock @3000px is ~60-70% parse (parse_pdf_content), not the fill. Parse = sequential content interpretation + per-Do draw_image (resample) + per-glyph Type1 charstring→outline→transformed-Shape building + render_soft_mask_group. Levers:

  1. Parallelize across images on image-heavy pages (page 8 = 1.27s, 34 JPEGs). draw_image runs sequentially per Do; the resample is parallel only within one image (parallel_for_rows), so 34 small images serialize. Defer image draws, resample+convert all in parallel, then composite in z-order. Risk: z-order / interleaving with vector content. Ceiling ~1.3-1.5× on page 8. Medium effort/risk.
  2. render_soft_mask_group is 10-14% on soft-mask pages. It renders each of ~17 groups to a fresh offscreen (capped kMaxSoftMaskGroupPixels=2 MP). Reuse a scratch buffer (member) to cut alloc/zero/page-fault churn; and/or lower the cap (masks are smooth gradients — near-identical OK) — but VERIFY fig-14 fade quality, it's the thing we just fixed. Low effort, low-moderate value.
  3. draw_image area-average resample: fixed-point reciprocal instead of 4 int divides/dst-pixel (±1 LSB; user OK'd near-identical). Small (~5%). Low effort.

B. Text rendering speed (deliberately NOT done — quality risk)

  • Glyph bitmap cache for Type1 (rasterize once, blit) would cut the 16% fill_polygon_aa on text pages, BUT the codebase deliberately bypasses bitmap caching for TJ body text (in_tj_text_draw_ gate) to keep subpixel positioning. Enabling it snaps glyphs to the pixel grid — a visible text change. Only do with explicit sign-off; verify text crops vs current. High value, real quality risk.

C. Embedded fonts (from older memory, not addressed this session)

  • Raw-CFF simple fonts (FontFile3) fail to load → substitute (math italic CM, Courier mono weight). cff_wrapper::wrap_cff_in_opentype emits OTTO with only the CFF table; stbtt/ttf_parse need head/hhea/maxp/hmtx(+cmap). Synthesize those + select glyphs by name via the CFF charset. See project_tex_math_glyph_names → "Biggest remaining opportunity". Larger cross-cutting change; affects other PDFs.

D. Encode / output

  • --fast-png (fpnge) is opt-in (files ~1.9× larger, lossless). If batch throughput matters more than size, consider a size-budgeted auto mode. Default stays miniz (smaller). Done as opt-in; revisit only on request.

E. Interruptible render — finer granularity / true cancel

  • Cancel is checked only at percent_step boundaries during parse; NOT mid- image-decode, mid-render_soft_mask_group, or during the parallel pre-decode barrier. For a single very heavy operator the budget can overshoot. Add checkpoints inside draw_image/render_soft_mask_group if needed. Low value unless a page has one dominant op.
  • WASM cancel is time-budget only (single-threaded JS can't set a flag mid-render). For true scroll-driven cancel, move rendering to a Web Worker + SharedArrayBuffer cancel flag (needs COOP/COEP headers on Pages) and have the budget callback also read the flag. nanopdf_last_render_interrupted() already exists. Larger; only if the time-budget UX proves insufficient.
  • --fast-png is not wired into the WASM viewer (CLI-only). Not needed for the viewer (it paints pixels, doesn't PNG-encode), but exposing it for the viewer's "Export/Save as PNG" paths could speed those. Low.

F. Throughput (batch --all, not single-page)

  • Cross-page parallelism (render N pages on N threads): ~3-4× for --all, but the global RenderCache LRU singleton causes cross-page output divergence on a few image/shading pages (documented gotcha) and needs per-thread Pdf. User previously chose single-page latency, so this was skipped. Revisit only if batch throughput becomes the goal; requires per-thread caches for determinism.

G. Validation breadth

  • Everything was tuned/verified against data/dcsdd.pdf only. Regression-render a wider corpus (other data/*.pdf, the math-font PDFs) to catch fonts/shadings the session's changes might affect — especially the dense shading-stop sampling and the transparency-group group_opacity. Cheap, worth doing before further changes.

How to resume / verify quickly

# build (rebuild BOTH or hit the stale-rasterize trap)
cmake --build build --target nanopdf -j8
cmake --build build-rasterize --target rasterize -j8
# perf baseline @3000px (TGA = render only; PNG includes encode)
for p in 1 7 8 5; do ./build-rasterize/rasterize data/dcsdd.pdf /tmp/p$p.tga -p $p -w 3000 --backend lightvg --format tga; done
# all pages -> compare_output/dcsdd_p300 (300 dpi)
for p in $(seq 1 14); do ./build-rasterize/rasterize data/dcsdd.pdf compare_output/dcsdd_p300/page_$(printf %02d $p).png -p $p --dpi 300; done
# pixel-diff a change vs prior binary: git stash, rebuild, render base, pop, rebuild, render new, compare in python
# tests
cd build && ctest --output-on-failure
# WASM (after src change): touch the .cc, rebuild, it auto-copies to examples/wasm/src/
source /mnt/nvme02/work/emsdk/emsdk_env.sh && cmake --build build_wasm -j8
# live viewer: vite dev on :5173 (examples/wasm), or the deployed Pages site

Perf method that worked: perf record -g --call-graph dwarf then perf report --no-children; but for WALL-CLOCK splits use direct std::chrono instrumentation — perf sums parallel-thread CPU and over-counts the parallel decode/resample.


Resuming prompt (paste to continue)

Continue optimizing/maintaining nanopdf rendering. State: main @ 1c6e089 is pushed and live on GitHub Pages. This session fixed dcsdd.pdf rendering (soft-mask gradation, transparency-group opacity, tx-fonts delimiters, Type1) and made rasterize ~2× faster at 3000px (soft-mask snapshot clip, parallel image decode/resample, tiled rasterization, opt-in fpnge) plus added interruptible rendering (bool progress callback + WASM time-budget render + viewer preview budgets). All pixel-identical, all tests green. See task.md "Remaining tasks" and the project_* memory notes. Next, I want to : (A) parallelize per-image draw on image-heavy pages (page 8); (B) reuse/shrink soft-mask group offscreen buffers; (C) load raw-CFF FontFile3 simple fonts (math italic / mono weight) instead of substituting; (D) regression-render a wider PDF corpus to validate this session's changes; (E) move WASM rendering to a Web Worker + SharedArrayBuffer for true scroll-cancel. Rebuild BOTH build/ and build-rasterize/ (stale-binary trap), verify pixel-identity vs the prior binary, and run ctest.