Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions guided-navigation/CHANGELOG.MD
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,17 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [1.1.0] – 2026-10-02

### Added

- `href` option for `makeGnd`/`parseMarkup`: every ref is resolved against the input's own href, e.g. `#p1` becomes `chapter.xhtml#p1`
- `DecodedTextref.href`: `decodeTextref` reads textrefs naming their resource (`chapter.xhtml#css(...)`, `chapter.xhtml#p1`) instead of ignoring them

### Changed

- `combineDomRangeTextrefs` returns `undefined` for references to two different resources

## [1.0.1] – 2026-10-02

### Added
Expand Down
8 changes: 6 additions & 2 deletions guided-navigation/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,8 @@ Malformed XHTML (an undeclared `epub:` prefix, an HTML entity like ` `…)

`makeGnd` returns `undefined` when the input has no navigable content, since a document's `guided` can't be empty. `parseMarkup` returns the objects without the document wrapper, and an empty array in that case.

Pass the input's own href as `{ href }` to resolve every ref (`textref`, `imgref`, `audioref`, `videoref`) against it: with `OEBPS/text/chapter.xhtml`, `#p1` becomes `OEBPS/text/chapter.xhtml#p1` and `../images/a.png` becomes `OEBPS/images/a.png`. Refs come out URL-encoded (`my chapter.xhtml` becomes `my%20chapter.xhtml`), so pass an encoded href to get it back unchanged.

Use `serialize()` on the result to get the JSON.

## Text references (`textrefs`)
Expand Down Expand Up @@ -90,13 +92,15 @@ If the node's text is too long, the polyfill only keeps the first few words and
Use `decodeTextref({ id, textref })` to read a `textref` back. It returns:

```typescript
{ cssSelector?, domRange?, text?, fragment? }
{ href?, cssSelector?, domRange?, text?, fragment? }
```

- `href` → the resource named before the fragment, when there is one (`chapter.xhtml#css(...)`)

- Exact match (one quote) → `text: { highlight, before?, after? }`
- Range (the `textStart`/`textEnd` case above) → `fragment` (the raw `:~:text=...` string)

It returns `undefined` if the `textref` isn't one of ours (e.g. it's a plain link's `href`). `combineDomRangeTextrefs(first, last)` joins two decoded `domRange` references into one spanning both.
It returns `undefined` if the `textref` is navigational rather than a reference to the node's own content, e.g. a link's `href`. `#css(...)` and `#domrange(...)` are always the node's own. A `#id` is the node's own only when it matches the node's `id`. `combineDomRangeTextrefs(first, last)` joins two decoded `domRange` references into one spanning both, as long as they're in the same resource.

The lower-level pieces are exported too: `encodeCssSelectorFragment(selector)`/`decodeCssSelectorFragment(textref)` for `#css(...)`, and `encodeDomRangeFragment(domRange)`/`decodeDomRangeFragment(textref)` for `#domrange(...)` (a `DomRangeJSON`: `{ start: { cssSelector, textNodeIndex, charOffset? }, end?: {...} }`, the RWPM shape behind `@readium/shared`'s `DomRange`).

Expand Down
2 changes: 1 addition & 1 deletion guided-navigation/package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@readium/guided-navigation",
"version": "1.0.1",
"version": "1.1.0",
"type": "module",
"description": "Guided Navigation document generation from HTML and XHTML content",
"author": "readium",
Expand Down
8 changes: 7 additions & 1 deletion guided-navigation/src/converter.ts
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@ import { ssmlAttrEscape, ssmlTextEscape, startsWithBindingPunct } from "@readium
import { type ObjBuilder, NavObject, isEmptyObj, finalizeToGuidedNavigationObject } from "./object.ts";
import { type GndMediaType, nodeLanguage, isInlineTag, sniffMediaType } from "./dom.ts";
import { encodeDomRangeFragment, encodeTextFragmentDirective } from "./textrefFragment.ts";
import { resolveRefs } from "./href.ts";
import { rootAnchorSelector, selectorForElement, textrefForSelector } from "./selectorGenerator.ts";
import { generateDomRange } from "./domRangeGenerator.ts";
import { textFragmentDirectiveFor } from "./textFragmentGenerator.ts";
Expand Down Expand Up @@ -70,6 +71,7 @@ export class Converter {
// See TextrefOptions.roles' leafTextRoleKeyword in options.ts.
leafTextEnabled = false;
docRoot: Document | null = null;
baseHref?: string;
// The live element selectors are climbed relative to — see
// selectorGenerator.ts's selectorForElement(). Only set when converting a
// live, already-rendered element (same caveat as domRangeEnabled).
Expand Down Expand Up @@ -161,7 +163,9 @@ export class Converter {
}

result(): GuidedNavigationObject[] {
return this.resultBuilders().map(finalizeToGuidedNavigationObject);
const builders = this.resultBuilders();
if (this.baseHref) builders.forEach((b) => resolveRefs(b, this.baseHref!));
return builders.map(finalizeToGuidedNavigationObject);
}

// An explicit-role descendant is content in its own right and skips
Expand Down Expand Up @@ -697,6 +701,7 @@ export function parseMarkup(
converter.domRangeEnabled = domRange;
converter.textFragmentEnabled = textFragment;
converter.docRoot = input.ownerDocument;
converter.baseHref = options?.href;
converter.selectorRoot = input;
converter.selectorRootAnchor = rootAnchorSelector(input);
converter.convert(input);
Expand All @@ -720,6 +725,7 @@ export function parseMarkup(
converter.leafTextEnabled = leafText;
converter.textFragmentEnabled = textFragment;
converter.docRoot = doc;
converter.baseHref = options?.href;
const body = doc.querySelector("body");
if (body && !BODY_TAG_RE.test(input)) {
converter.convertChildren(body);
Expand Down
26 changes: 26 additions & 0 deletions guided-navigation/src/href.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
import type { ObjBuilder } from "./object.ts";

const SCHEME_RE = /^[a-z][a-z\d+.-]*:/i;
// Placeholder origin so URL can resolve relative paths; stripped from the result.
const BASE = "https://gnd.invalid/";

// Resolves a ref found in the resource at `href` against it, keeping the result relative to the same base as `href`.
export function resolveRef(ref: string, href: string): string {
if (SCHEME_RE.test(ref)) return ref;
// The fragment is kept verbatim: URL parsing would re-encode textref fragments.
const hashIndex = ref.indexOf("#");
const path = hashIndex === -1 ? ref : ref.slice(0, hashIndex);
const fragment = hashIndex === -1 ? "" : ref.slice(hashIndex);
if (SCHEME_RE.test(href)) return new URL(path, href).href + fragment;
if (href.startsWith("//")) return new URL(path, "https:" + href).href.slice("https:".length) + fragment;
const resolved = new URL(path, BASE + href.replace(/^\/+/, "")).href;
return resolved.startsWith(BASE) ? resolved.slice(BASE.length) + fragment : ref;
}

export function resolveRefs(o: ObjBuilder, href: string): void {
if (o.textref) o.textref = resolveRef(o.textref, href);
if (o.imgref) o.imgref = resolveRef(o.imgref, href);
if (o.audioref) o.audioref = resolveRef(o.audioref, href);
if (o.videoref) o.videoref = resolveRef(o.videoref, href);
o.children?.forEach((c) => resolveRefs(c, href));
}
3 changes: 3 additions & 0 deletions guided-navigation/src/options.ts
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,9 @@ export interface GndGenerationOptions {
// Re-parse string input as text/html when XHTML parsing fails, instead of
// throwing.
htmlFallback?: boolean;
// The input's own href: every ref (textref, imgref, audioref, videoref) is
// resolved against it, e.g. "#p1" becomes "chapter.xhtml#p1".
href?: string;
}

function normalizeRoles(opt: boolean | GndRole[] | undefined): ((roles: GndRole[]) => boolean) | null {
Expand Down
40 changes: 24 additions & 16 deletions guided-navigation/src/textrefFragment.ts
Original file line number Diff line number Diff line change
Expand Up @@ -108,6 +108,8 @@ export function decodeTextFragmentDirective(textref: string | undefined): TextFr
}

export interface DecodedTextref {
// The resource the reference points into, when the textref names one.
href?: string;
cssSelector?: string;
domRange?: DomRangeJSON;
text?: { highlight?: string; before?: string; after?: string };
Expand All @@ -120,9 +122,13 @@ export interface DecodedTextref {
// `domRange` to combine — a bare selector has no `textNodeIndex` to build one from.
export function combineDomRangeTextrefs(first: DecodedTextref, last: DecodedTextref): DecodedTextref | undefined {
if (!first.domRange || !last.domRange) return undefined;
// A DOM range can't span two documents.
if (first.href !== last.href) return undefined;
const domRange: DomRangeJSON = { start: first.domRange.start, end: last.domRange.end ?? last.domRange.start };
if (first.domRange.container !== undefined) domRange.container = first.domRange.container;
return { domRange, cssSelector: domRange.container ?? domRange.start.cssSelector };
const combined: DecodedTextref = { domRange, cssSelector: domRange.container ?? domRange.start.cssSelector };
if (first.href !== undefined) combined.href = first.href;
return combined;
}

function decodeIdFragment(base: string): string | undefined {
Expand All @@ -134,33 +140,34 @@ function decodeIdFragment(base: string): string | undefined {
}
}

// Decodes a node's own generated textref, distinguishing it from an
// unrelated navigational textref (link href, pagebreak/noteref reference)
// that happens to also start with "#" — those are never wrapped in
// "#css(...)"/"#domrange(...)", and a bare "#id" is only trusted as a
// self-reference when it matches this same node's own id (the shape
// Converter.applyTextref produces), not an id belonging elsewhere. A
// ":~:text=..." suffix is independent of all that — it can accompany any of
// the above, or stand alone — so it's decoded separately from the rest of
// the string and merged into the result.
// Decodes a reference to a node's own content, distinguishing it from a
// navigational textref (link href, noteref target). A textref is
// "[href]#fragment", the href naming the resource when present (e.g.
// "chapter.xhtml#css(...)"). "#css(...)"/"#domrange(...)" fragments are
// always self-references; a bare "#id" only is when it matches the node's
// own id. A ":~:text=..." suffix can accompany any of the above, or stand
// alone, so it's decoded separately and merged into the result.
export function decodeTextref(node: { id?: string; textref?: string } | undefined): DecodedTextref | undefined {
const textref = node?.textref;
if (!textref) return undefined;

const markIndex = textref.indexOf(":~:");
const base = markIndex === -1 ? textref : textref.slice(0, markIndex);
const hashIndex = base.indexOf("#");
const href = hashIndex === -1 ? base : base.slice(0, hashIndex);
const fragment = hashIndex === -1 ? "" : base.slice(hashIndex);

let cssSelector: string | undefined;
let domRange: DomRangeJSON | undefined;
const baseDomRange = decodeDomRangeFragment(base);
if (baseDomRange) {
domRange = baseDomRange;
cssSelector = baseDomRange.container ?? baseDomRange.start.cssSelector;
const fragmentDomRange = decodeDomRangeFragment(fragment);
if (fragmentDomRange) {
domRange = fragmentDomRange;
cssSelector = fragmentDomRange.container ?? fragmentDomRange.start.cssSelector;
} else {
const decoded = decodeCssSelectorFragment(base);
const decoded = decodeCssSelectorFragment(fragment);
if (decoded !== undefined) {
cssSelector = decoded;
} else if (node?.id && decodeIdFragment(base) === node.id) {
} else if (node.id && decodeIdFragment(fragment) === node.id) {
cssSelector = `#${CSS.escape(node.id)}`;
}
}
Expand All @@ -177,6 +184,7 @@ export function decodeTextref(node: { id?: string; textref?: string } | undefine
}

const result: DecodedTextref = {};
if (href !== "") result.href = href;
if (cssSelector !== undefined) result.cssSelector = cssSelector;
if (domRange !== undefined) result.domRange = domRange;
if (directive?.textEnd !== undefined) {
Expand Down
75 changes: 75 additions & 0 deletions guided-navigation/test/textref.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,9 @@ import {
encodeTextFragmentDirective,
decodeTextFragmentDirective,
decodeTextref,
combineDomRangeTextrefs,
} from "../src/textrefFragment.ts";
import { resolveRef } from "../src/href.ts";

test("textrefs option is off by default — no textref is generated", () => {
const [result] = parseMarkup("<p>Hello.</p>");
Expand Down Expand Up @@ -423,3 +425,76 @@ test("decodeTextref carries a textStart/textEnd range as fragment, not text.high
fragment: encodeTextFragmentDirective(directive),
});
});

test("decodeTextref splits an href from a self-referencing fragment", () => {
expect(decodeTextref({ textref: `page3.xhtml${encodeCssSelectorFragment("p")}` })).toEqual({ href: "page3.xhtml", cssSelector: "p" });
const domRange = { start: { cssSelector: "p", textNodeIndex: 0 } };
expect(decodeTextref({ textref: `page3.xhtml${encodeDomRangeFragment(domRange)}` })).toEqual({ href: "page3.xhtml", cssSelector: "p", domRange });
expect(decodeTextref({ id: "p12", textref: "chapter.xhtml#p12" })).toEqual({ href: "chapter.xhtml", cssSelector: "#p12" });
});

test("decodeTextref ignores href-qualified navigational textrefs", () => {
expect(decodeTextref({ textref: "endnotes.xhtml#note2" })).toBe(undefined);
expect(decodeTextref({ id: "p1", textref: "chapter.xhtml#p2" })).toBe(undefined);
expect(decodeTextref({ textref: "chapter.xhtml" })).toBe(undefined);
});

test("decodeTextref ignores a link's href even when the link has a role", () => {
const [landmarks] = parseMarkup(
'<ol xmlns="http://www.w3.org/1999/xhtml" xmlns:epub="http://www.idpf.org/2007/ops"><li><a epub:type="toc" href="#toc">Contents</a></li><li><a epub:type="bodymatter" href="chapter.xhtml">Start</a></li></ol>',
"application/xhtml+xml",
);
const links = landmarks.children?.flatMap((li) => li.children ?? [li]) ?? [];
expect(links.map((l) => l.textref)).toEqual(["#toc", "chapter.xhtml"]);
for (const link of links) expect(decodeTextref(link)).toBe(undefined);
});

test("decodeTextref keeps the href alongside a text-fragment directive", () => {
const textref = `page3.xhtml${encodeCssSelectorFragment("p")}${encodeTextFragmentDirective({ textStart: "Hello." })}`;
expect(decodeTextref({ textref })).toEqual({ href: "page3.xhtml", cssSelector: "p", text: { highlight: "Hello." } });
});

test("combineDomRangeTextrefs refuses to span two resources", () => {
const start = { cssSelector: "p", textNodeIndex: 0 };
expect(combineDomRangeTextrefs({ href: "p1.xhtml", domRange: { start } }, { href: "p2.xhtml", domRange: { start } })).toBe(undefined);
expect(combineDomRangeTextrefs({ href: "p1.xhtml", domRange: { start } }, { href: "p1.xhtml", domRange: { start } })?.href).toBe("p1.xhtml");
});

test("resolveRef resolves fragments, relative paths and keeps absolute URLs", () => {
expect(resolveRef("#p1", "OEBPS/chapter.xhtml")).toBe("OEBPS/chapter.xhtml#p1");
expect(resolveRef("../images/a.png", "OEBPS/text/chapter.xhtml")).toBe("OEBPS/images/a.png");
expect(resolveRef("notes.xhtml#n1", "OEBPS/chapter.xhtml")).toBe("OEBPS/notes.xhtml#n1");
expect(resolveRef("https://example.com/x", "OEBPS/chapter.xhtml")).toBe("https://example.com/x");
expect(resolveRef("a.png", "https://example.com/pub/chapter.xhtml")).toBe("https://example.com/pub/a.png");
expect(resolveRef("#p1", "//cdn.example/pub/chapter.xhtml")).toBe("//cdn.example/pub/chapter.xhtml#p1");
expect(resolveRef("../a.png", "//cdn.example/pub/chapter.xhtml")).toBe("//cdn.example/a.png");
});

test("resolveRef normalizes the href the same way for #fragment refs and path refs", () => {
for (const href of ["/OEBPS/chapter.xhtml", "OEBPS/./text/../chapter.xhtml"]) {
expect(resolveRef("#p1", href)).toBe("OEBPS/chapter.xhtml#p1");
expect(resolveRef("chapter.xhtml#p2", href)).toBe("OEBPS/chapter.xhtml#p2");
}
expect(resolveRef("#p1", "OEBPS/my chapter.xhtml")).toBe("OEBPS/my%20chapter.xhtml#p1");
expect(resolveRef("my chapter.xhtml#p1", "OEBPS/my chapter.xhtml")).toBe("OEBPS/my%20chapter.xhtml#p1");
expect(resolveRef("#p1", "OEBPS/my%20chapter.xhtml")).toBe("OEBPS/my%20chapter.xhtml#p1");
const css = encodeCssSelectorFragment("p > span");
expect(resolveRef(css, "OEBPS/chapter.xhtml")).toBe(`OEBPS/chapter.xhtml${css}`);
});

test("parseMarkup with href qualifies every ref, and the textrefs decode back with that href", () => {
const nodes = parseMarkup(
'<p id="p1">Hello <a href="notes.xhtml#n1">note</a>.</p><p>World.</p><img src="a.png" alt="A"/>',
"text/html",
{ textrefs: true, href: "OEBPS/chapter.xhtml" },
);
expect(nodes[0].textref).toBe("OEBPS/chapter.xhtml#p1");
expect(decodeTextref(nodes[0])).toEqual({ href: "OEBPS/chapter.xhtml", cssSelector: "#p1" });
const link = nodes[0].children?.find((c) => c.textref?.includes("notes"));
expect(link?.textref).toBe("OEBPS/notes.xhtml#n1");
expect(decodeTextref(link)).toBe(undefined);
const second = decodeTextref(nodes[1]);
expect(second?.href).toBe("OEBPS/chapter.xhtml");
expect(second?.cssSelector).toBeTruthy();
expect(nodes[2].imgref).toBe("OEBPS/a.png");
});
Loading