Skip to content

Use ASCII case folding for HTML element and attribute names - #200

Open
ac1982 wants to merge 1 commit into
scrapy:masterfrom
ac1982:fix/ascii-html-name-folding
Open

ac1982 wants to merge 1 commit into
scrapy:masterfrom
ac1982:fix/ascii-html-name-folding

Conversation

@ac1982

@ac1982 ac1982 commented Oct 10, 2026

Copy link
Copy Markdown

HTMLTranslator currently Unicode-lowercases element and attribute names. Given <x-Ä id="upper"></x-Ä><x-ä id="lower"></x-ä>, selecting x-Ä incorrectly returns lower. [DATA-Ä] likewise selects a data-ä attribute instead of data-Ä.

Use the existing ascii_lower helper for HTML name normalization, consistent with HTML/CSS ASCII case-matching rules. ASCII portions still fold (X-Ä matches x-Ä), while non-ASCII characters retain their identity. XML/XHTML and the attribute-value extension point are unchanged.

Two regressions fail against master and pass with the fix; they execute the generated XPath against actual lxml-parsed HTML and cover ordinary, mixed-ASCII-case, and wildcard-namespace selectors. This is unrelated to #50 (Python 2 byte-string input), #189/#191 (tokenization), and the recent attribute-value i flag support.

Validation on Python 3.14.4 / Linux arm64:

  • Full tests including documentation examples: 29 passed; XPath module 100% coverage, total 99%.
  • Ruff lint/format, mypy (5 source files), pylint, complete pre-commit hooks, Sphinx build with -W, sdist build / twine check, and git diff --check passed.
  • Other Python versions, PyPy, macOS and Windows were not run locally.

AI assistance: OpenAI Codex investigated, implemented, and ran these checks.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant