Repository navigation
Conversation
`next_token()` pushed every HTML span and void block onto two parallel stacks and popped it again on the next call, and deferred stack updates behind a string-valued pending operation. Store only the offsets of open blocks. A paused HTML span or void block adds one level of depth in `get_depth()` and one breadcrumb in `get_breadcrumbs()`. Openers push and closers pop when the processor reaches them. Block type lengths are computed when breadcrumbs are requested, because a block type is always followed by whitespace. Scanning a document with `next_token()` takes 13% to 15% less time. Adds tests for breadcrumbs and depth on HTML spans, void blocks, and blocks left open at the end of a document. See #66138.
`next_delimiter()` called `next_token()` once for each HTML span and again for the delimiter after it, even though only the wildcard and freeform block types can match an HTML span. For every other search, `next_token()` now continues past the HTML span in the same call. The processor visits the same delimiters and stops in the same state as before, including when the document ends in an HTML span or in a partial delimiter. Visiting every delimiter and reading its type, block type, attributes and span takes 9% to 13% less time over block posts, theme templates and Gutenberg's block fixtures. `WP_Block_Parser` on top of the processor (WordPress#13695) takes 6% to 10% less. Adds a test which compares `next_delimiter()` against filtering `next_token()` for several block types and document endings. See #66138.
`strpos( $text, '<!--' )` calls `memchr()` for `<`, which stops at every HTML tag. Search for `!--` instead and check the byte before it. In HTML `!` is far less common than `<`: the classic post used for measurement has 4,510 `<` and 44 `!`. Scanning a 300 KB classic post without block delimiters takes 80% less time (27 µs to 5.6 µs), the same as the regex in `WP_Block_Parser`. Block content shows no change. Adds tests for `!--` sequences outside of comment openers. See #66138.
For every delimiter candidate, `next_token()` called `find_html_comment_end()`, which first checks for a comment made only of dashes and then searches for `--` from the start of the comment. A delimiter candidate begins with `<!--` and whitespace, and its block type is followed by whitespace, so its comment cannot be a dash-only comment and its closer cannot start before the JSON span. Search for the closer from there in `next_token()`. When there is no closer the input is incomplete, as before. Scanning block documents with `next_token()` takes 10% to 12% less time, and `WP_Block_Parser` on top of the processor (WordPress#13695) 5% to 7% less. Adds tests for dash runs inside block types. See #66138.
`next_token()`, `next_delimiter()`, `next_block()` and `extract_full_block_and_advance()` change the processor's state, but PHPStan assumed repeated calls return the same value. Inside `extract_full_block_and_advance()` it therefore reported the checks for HTML spans after `next_token()` as always false, which the baseline ignored. Mark the four methods `@phpstan-impure`, as the Tag Processor does for `next_tag()`, and remove the two baseline entries. See #66138.
`get_block_type()` compared the state against three constants, called `is_html()`, and called `normalize_block_type()`, which searches the block type for `/`. Only a matched delimiter has a block type, and the processor already knows whether a namespace is present: the name starts where the namespace would have. Check for the matched state once and add the implicit `core/` namespace when the name and namespace offsets are equal. `get_printable_block_type()` now handles HTML spans and defers to `get_block_type()`. Methods inside the class compare the state against `HTML_SPAN` instead of calling `is_html()`. Reading the type, block type, attributes and span of every delimiter takes 6% to 8% less time, and `WP_Block_Parser` on top of the processor (WordPress#13695) 4% to 5% less. Adds tests that no block type is reported before scanning, after the end of the document, or after an error. See #66138.
`extract_full_block_and_advance()` called `get_depth()` for every token, `get_html_content()` for every HTML span, and `opens_block()` for every delimiter. Compute the same values in place. Extracting every top-level block takes 3% to 9% less time. See #66138.
`next_delimiter( $block_type )` called `is_block_type()` for every delimiter, which calls `are_equal_block_types()`. That scans both block types for `/` and builds a `core/` string whenever exactly one of them has a namespace. Normalize the searched block type once per call. A delimiter whose block type has no namespace then matches only the part after `core/`, and every other delimiter matches the full type. Both comparisons check the length first. `next_block( $block_type )` takes 13% to 21% less time. Adds explicit namespaces and search types which cannot match to the test comparing `next_delimiter()` with `next_token()`. See #66138.
The closer search used `strpos( $text, '--' )`, which stops at every dash in the JSON attributes, and then `strspn()` over the dash run. Search for `>` instead. No closer can end before the first `>` after the start of the JSON span, so that `>` closes the comment if `--` or `--!` precedes it, and otherwise the search continues after it. The block editor escapes `>` in attributes as `>`, so the first `>` is usually the closer. Scanning theme templates with `next_token()` takes 3% less time; other documents change by 1% to 2%. Adds tests for `>` inside of the JSON attributes. See #66138.
Reproduce the differential trace and the timingsThe script compares mkdir block-processor-verify && cd block-processor-verify
git clone --depth=1 https://github.com/WordPress/gutenberg.git
git clone --filter=blob:none https://github.com/WordPress/wordpress-develop.git
git -C wordpress-develop fetch origin pull/14117/head:block-processor-perf pull/13695/head:block-parser-13695
git -C wordpress-develop worktree add --detach ../trunk "$(git -C wordpress-develop merge-base origin/trunk block-processor-perf)"
git -C wordpress-develop worktree add ../pr block-processor-perf
git -C wordpress-develop show block-parser-13695:src/wp-includes/class-wp-block-parser.php > parser-13695.php
# Save the script below as block-processor-verify.php, then:
php -d opcache.enable_cli=1 -d opcache.file_update_protection=0 block-processor-verify.php trunk pr gutenberg --parser=parser-13695.php
<?php
/**
* Compares WP_Block_Processor in two wordpress-develop checkouts.
*
* 1. Differential trace: every public accessor after every step, in 14 traversal
* modes, over the corpus, edge cases, truncations and mutated documents.
* 2. Benchmark: median time per pass over each document set. Every round runs each
* variant and work in its own PHP process, in random order, with this process's
* opcache and JIT settings. Variants in one process would share JIT traces.
*
* Usage:
* php -d opcache.enable_cli=1 -d opcache.file_update_protection=0 block-processor-verify.php \
* <trunk-checkout> <pr-checkout> <gutenberg-checkout> \
* [--parser=<file>] [--rounds=15] [--mutations=20000] [--no-diff] [--no-bench]
*
* --parser: `src/wp-includes/class-wp-block-parser.php` from wordpress-develop#13695.
* Also times that parser on each processor against the regex parser in <trunk-checkout>.
*
* Without opcache.file_update_protection=0, opcache does not cache the variant files
* this script writes, because they are less than two seconds old.
*/
// Polyfills for PHP 7.4 to 8.3, as in wp-includes/compat.php.
if ( ! function_exists( 'str_starts_with' ) ) {
function str_starts_with( $h, $n ) { return 0 === strncmp( $h, $n, strlen( $n ) ); }
function str_ends_with( $h, $n ) { return '' === $n || ( strlen( $h ) >= strlen( $n ) && 0 === substr_compare( $h, $n, -strlen( $n ) ) ); }
function str_contains( $h, $n ) { return false !== strpos( $h, $n ); }
}
if ( ! function_exists( 'array_any' ) ) {
function array_any( array $array, callable $callback ) { foreach ( $array as $k => $v ) { if ( $callback( $v, $k ) ) { return true; } } return false; }
}
$positional = array();
$parser = null;
$rounds = 15;
$mutations = 20000;
$run_diff = true;
$run_bench = true;
$child = null;
$child_work = null;
foreach ( array_slice( $argv, 1 ) as $arg ) {
if ( str_starts_with( $arg, '--child=' ) ) {
// Internal: time one variant and one work, then print the results as JSON.
$child = substr( $arg, 8 );
} elseif ( str_starts_with( $arg, '--work=' ) ) {
$child_work = substr( $arg, 7 );
} elseif ( str_starts_with( $arg, '--parser=' ) ) {
$parser = substr( $arg, 9 );
} elseif ( str_starts_with( $arg, '--rounds=' ) ) {
$rounds = max( 9, (int) substr( $arg, 9 ) );
} elseif ( str_starts_with( $arg, '--mutations=' ) ) {
$mutations = (int) substr( $arg, 12 );
} elseif ( '--no-diff' === $arg ) {
$run_diff = false;
} elseif ( '--no-bench' === $arg ) {
$run_bench = false;
} else {
$positional[] = rtrim( $arg, '/' );
}
}
if ( 3 !== count( $positional ) ) {
fwrite( STDERR, "Usage: php -d opcache.enable_cli=1 -d opcache.file_update_protection=0 {$argv[0]} <trunk-checkout> <pr-checkout> <gutenberg-checkout> [--parser=<file>] [--rounds=N] [--mutations=N] [--no-diff] [--no-bench]\n" );
exit( 2 );
}
list( $trunk, $pr, $gutenberg ) = $positional;
foreach ( array( "$trunk/src/wp-includes/class-wp-block-processor.php", "$pr/src/wp-includes/class-wp-block-processor.php", "$gutenberg/test/performance/assets/large-post.html" ) as $required ) {
if ( ! is_file( $required ) ) {
fwrite( STDERR, "Missing $required\n" );
exit( 2 );
}
}
$opcache = function_exists( 'opcache_is_script_cached' ) && ini_get( 'opcache.enable_cli' );
$jit = $opcache && ( opcache_get_status( false )['jit']['on'] ?? false );
if ( null === $child ) {
printf( "PHP %s, opcache %s, JIT %s\n", PHP_VERSION, $opcache ? 'on' : 'off', $jit ? 'on' : 'off' );
if ( ! $opcache ) {
echo "Warning: opcache is off, so times differ from a production site.\n";
}
if ( null !== $parser && str_contains( file_get_contents( "$trunk/src/wp-includes/class-wp-block-parser.php" ), 'WP_Block_Processor' ) ) {
fwrite( STDERR, "The parser in $trunk uses WP_Block_Processor, so it is not the regex parser.\n" );
exit( 2 );
}
} else {
$run_diff = false;
}
require "$trunk/src/wp-includes/html-api/class-wp-html-span.php";
require "$trunk/src/wp-includes/class-wp-block-parser-block.php";
require "$trunk/src/wp-includes/class-wp-block-parser-frame.php";
/*
* Each variant is loaded from its own file path under renamed classes,
* because opcache caches by path.
*/
$build = sys_get_temp_dir() . '/block-processor-verify-' . getmypid();
mkdir( $build );
function load_renamed( string $build, string $file, string $name, array $renames ): void {
global $opcache;
$source = file_get_contents( $file );
foreach ( $renames as $from => $to ) {
$source = preg_replace( "~\\b{$from}\\b~", $to, $source );
}
// The parser files require the block and frame classes, loaded above.
$source = preg_replace( '~^require_once __DIR__ .*$~m', '', $source );
$path = "$build/$name.php";
file_put_contents( $path, $source );
require $path;
if ( $opcache && ! opcache_is_script_cached( $path ) ) {
fwrite( STDERR, "Warning: opcache did not cache $name; set -d opcache.file_update_protection=0.\n" );
}
unlink( $path );
}
$processors = array(
'trunk' => 'WP_Block_Processor_Trunk',
'PR' => 'WP_Block_Processor_PR',
);
$sources = array(
'trunk' => "$trunk/src/wp-includes/class-wp-block-processor.php",
'PR' => "$pr/src/wp-includes/class-wp-block-processor.php",
);
if ( $run_diff ) {
foreach ( $processors as $tag => $class ) {
load_renamed( $build, $sources[ $tag ], "processor-$tag", array( 'WP_Block_Processor' => $class ) );
}
}
// A child loads one processor and, to time `parse`, the #13695 parser on it; or the regex parser.
if ( null !== $child && isset( $processors[ $child ] ) ) {
load_renamed( $build, $sources[ $child ], "processor-$child", array( 'WP_Block_Processor' => $processors[ $child ] ) );
if ( 'parse' === $child_work ) {
load_renamed( $build, $parser, "parser-13695-$child", array( 'WP_Block_Parser' => "WP_Block_Parser_13695_$child", 'WP_Block_Processor' => $processors[ $child ] ) );
}
} elseif ( 'regex' === $child ) {
load_renamed( $build, "$trunk/src/wp-includes/class-wp-block-parser.php", 'parser-regex', array( 'WP_Block_Parser' => 'WP_Block_Parser_Regex' ) );
}
rmdir( $build );
/*
* Corpus.
*/
$corpus = array();
// Gutenberg's editor performance fixtures.
$corpus['large-post'] = array( file_get_contents( "$gutenberg/test/performance/assets/large-post.html" ) );
$corpus['small-post'] = array( file_get_contents( "$gutenberg/test/performance/assets/small-post-with-containers.html" ) );
// Gutenberg's block fixtures: one document per block type and variation.
$docs = array();
foreach ( glob( "$gutenberg/test/integration/fixtures/blocks/*.html" ) as $f ) {
if ( ! str_ends_with( $f, '.serialized.html' ) ) {
$docs[] = file_get_contents( $f );
}
}
$corpus['gb-fixtures'] = $docs;
// Core's block theme templates, parts and patterns.
$docs = array();
foreach ( array( 'twentytwentytwo', 'twentytwentythree', 'twentytwentyfour', 'twentytwentyfive' ) as $theme ) {
$dir = "$trunk/src/wp-content/themes/$theme";
if ( ! is_dir( $dir ) ) {
continue;
}
foreach ( new RecursiveIteratorIterator( new RecursiveDirectoryIterator( $dir, FilesystemIterator::SKIP_DOTS ) ) as $f ) {
$path = $f->getPathname();
if ( str_ends_with( $path, '.html' ) || ( str_contains( $path, '/patterns/' ) && str_ends_with( $path, '.php' ) ) ) {
$docs[] = file_get_contents( $path );
}
}
}
sort( $docs );
$corpus['theme-templates'] = $docs;
// A classic post: the large post without block delimiters, plus a more tag.
$classic = preg_replace( '~<!-- /?wp:[^>]*-->\n?~', '', $corpus['large-post'][0] );
$at = strpos( $classic, '</p>' ) + 4;
$classic = substr( $classic, 0, $at ) . "\n<!--more-->\n" . substr( $classic, $at );
$corpus['classic-post'] = array( $classic );
if ( null === $child ) {
echo "\nCorpus\n";
foreach ( $corpus as $set => $docs ) {
$bytes = array_sum( array_map( 'strlen', $docs ) );
$delims = array_sum( array_map( function ( $d ) { return preg_match_all( '~<!-- /?wp:~', $d ); }, $docs ) );
printf( " %-16s %4d docs %8d bytes %6d delimiters\n", $set, count( $docs ), $bytes, $delims );
}
}
/*
* Differential trace.
*/
function snapshot( $p ): array {
$span = $p->get_span();
$attrs = $p->allocate_and_return_parsed_attributes();
return array(
'span' => $span ? array( $span->start, $span->length ) : null,
'html' => $p->is_html(),
'nwhtml' => $p->is_non_whitespace_html(),
'content' => $p->get_html_content(),
'dtype' => $p->get_delimiter_type(),
'btype' => $p->get_block_type(),
'ptype' => $p->get_printable_block_type(),
'depth' => $p->get_depth(),
'crumbs' => $p->get_breadcrumbs(),
'opens' => $p->opens_block(),
'opens_p' => $p->opens_block( 'core/paragraph', 'group' ),
'is_any' => $p->is_block_type( '*' ),
'is_ff' => $p->is_block_type( 'core/freeform' ),
'is_ff2' => $p->is_block_type( 'freeform' ),
'is_p' => $p->is_block_type( 'paragraph' ),
'is_cp' => $p->is_block_type( 'core/paragraph' ),
'closing' => $p->has_closing_flag(),
'attrs' => $attrs,
'json_err' => $p->get_last_json_error(),
'error' => $p->get_last_error(),
);
}
function trace( string $class, string $doc, string $mode ): array {
$p = new $class( $doc );
$out = array( snapshot( $p ) );
for ( $i = 0; $i < 100000; $i++ ) {
switch ( $mode ) {
case 'token':
$ok = $p->next_token();
break;
case 'delim':
$ok = $p->next_delimiter();
break;
case 'delim*':
$ok = $p->next_delimiter( '*' );
break;
case 'delim-ff':
$ok = $p->next_delimiter( 'core/freeform' );
break;
case 'delim-p':
$ok = $p->next_delimiter( 'paragraph' );
break;
case 'delim-ns':
$ok = $p->next_delimiter( 'my/paragraph' );
break;
case 'delim-core-ns':
$ok = $p->next_block( 'core/paragraph' );
break;
case 'delim-odd':
// Search types which can never match.
$odd = array( '', 'core/', '/paragraph', 'a/b/c', 'Paragraph', 'core//paragraph' );
$ok = $p->next_delimiter( $odd[ $i % count( $odd ) ] );
break;
case 'block':
$ok = $p->next_block();
break;
case 'block*':
$ok = $p->next_block( '*' );
break;
case 'block-group':
$ok = $p->next_block( 'core/group' );
break;
case 'extract':
$ok = $p->next_block( '*' );
if ( $ok ) {
$out[] = $p->extract_full_block_and_advance();
}
break;
case 'mixed':
// Interleave traversal methods, starting from different token kinds.
$methods = array( 'next_token', 'next_delimiter', 'next_token', 'next_token', 'next_block', 'next_delimiter', 'next_token' );
$method = $methods[ $i % count( $methods ) ];
$ok = 1 === $i % 5 ? $p->next_delimiter( 'list-item' ) : $p->$method();
break;
case 'extract-inner':
// Extract every block at depth > 1, exercising mid-stream extraction.
$ok = $p->next_token();
if ( $ok && $p->get_depth() > 1 && $p->opens_block() ) {
$out[] = $p->extract_full_block_and_advance();
}
break;
}
$out[] = $ok;
$out[] = snapshot( $p );
if ( ! $ok ) {
// Once false, it should stay false.
$out[] = $p->next_token();
$out[] = snapshot( $p );
break;
}
}
return $out;
}
if ( $run_diff ) {
$docs = array();
foreach ( $corpus as $set => $set_docs ) {
foreach ( $set_docs as $i => $doc ) {
$docs[ "$set#$i" ] = $doc;
}
}
// Hand-written edge cases.
$edge = array(
'', ' ', "\n", 'x', '<', '<!', '<!-', '<!--', '<!-- ', '<!-- w', '<!-- wp', '<!-- wp:', '<!-- wp:a', '<!-- wp:a ', '<!-- wp:a -', '<!-- wp:a --', '<!-- wp:a -->',
'<!-- wp:a /-->', '<!-- /wp:a -->', '<!-- /wp:a /-->', '<!-- wp:a {} -->', '<!-- wp:a {}/-->', '<!-- wp:a {} /-->', '<!-- wp:a {"x":1}-->',
'<!-- wp:a --!>', '<!-- wp:a -- -->', '<!-- wp:a--b -->', '<!-- wp:a/b/c -->', '<!-- wp:A -->', '<!-- wp:1a -->', '<!-- wp:a/1 -->', '<!-- wp:a/ -->',
'<!----><!-- wp:a /-->', '<!---><!-- wp:a /-->', '<!--><!-- wp:a /-->', '<!--x--><!-- wp:a /-->', '<!-- x --><!-- wp:a /-->', '<!-- wp:a {"--":1} /-->',
'<!-- wp:a {"a":"-->"} /-->', '<!-- wp:a {"a":"}"} } /-->', '<!-- wp:a {"a":1}} /-->', '<!-- wp:a {"a":1} x /-->', '<!-- wp:a x /-->',
"<!--\twp:a\t/-->", "<!--\fwp:a\f-->", "<!--\rwp:a\r-->", "<!--\vwp:a\v-->", "<!--\nwp:a {}\n/-->",
'<!-- wp:a --><!-- wp:b --><!-- /wp:b --><!-- /wp:a -->', '<!-- wp:a -->x<!-- wp:b -->y<!-- /wp:b -->z<!-- /wp:a -->w',
'<!-- /wp:a -->', 'x<!-- /wp:a -->y', '<!-- wp:a -->', '<!-- wp:a -->x', 'x<!-- wp:a -->', '<!-- wp:a --><!-- /wp:b -->', '<!-- /wp:a --><!-- /wp:a --><!-- wp:b /-->',
'<!-- wp:a /--><!-- wp:a /-->', 'x<!-- wp:a /-->y<!-- wp:a /-->z', '<!-- wp:a {"x":1} /-->', '<!-- wp:a {"x":} /-->', '<!-- /wp:a {"x":1} -->',
'<!-- wp:a [1] /-->', '<!-- wp:a {"x":"\xff"} /-->', '<!-- wp:a {"x":1} -->x<!-- /wp:a -->', 'text<!--', 'text<!-', 'text<!', 'text<',
'<!-- wp:core/paragraph --><!-- /wp:paragraph -->', '<!-- wp:paragraph --><!-- /wp:core/paragraph -->', '<!-- wp:my/paragraph /-->',
'<!-- wp:a {"x":1} /--><!-- wp:b /-->', '<!-- wp:a -->' . str_repeat( '<!-- wp:b -->', 5 ) . 'x' . str_repeat( '<!-- /wp:b -->', 5 ) . '<!-- /wp:a -->',
'<!-- wp:freeform /-->', '<!-- wp:core/freeform -->x<!-- /wp:core/freeform -->', "<!-- wp:a -->\n\n<!-- /wp:a -->\n\n",
'<!-- wp:a {"x":"--"} /-->', '<!-- wp:a {"x":"<!--"} /-->', '<!-- wp:a {"x":"/-->"} /-->', '<!-- wp:a -->x<!-- /wp:a', '<!-- wp:a -->x<!-- /wp:a -',
'<!-- wp:a ---->', '<!-- wp:a --->', '<!-- wp:a /--->', '<!-- wp:a {} --->', '<!-- wp:a {} ---!>', '<!-- wp:a {}--!>',
);
foreach ( $edge as $i => $doc ) {
$docs[ "edge#$i" ] = $doc;
// Also embed each edge case between real content.
$docs[ "edge-ctx#$i" ] = "<p>Lead</p>\n<!-- wp:paragraph -->\n<p>x</p>\n<!-- /wp:paragraph -->\n{$doc}\n<!-- wp:group --><div><!-- wp:image {\"id\":1} /--></div><!-- /wp:group -->\ntail";
}
// The small post truncated at every byte.
$sample = $corpus['small-post'][0];
for ( $i = 0; $i <= strlen( $sample ); $i++ ) {
$docs[ "trunc#$i" ] = substr( $sample, 0, $i );
}
// Deterministic mutations of short real documents.
mt_srand( 1234 );
$alphabet = array( '<', '!', '-', '/', 'w', 'p', ':', '{', '}', '"', ' ', "\n", "\t", '>', 'a', 'z', '0', '_', 'core/', '<!--', '-->', '/-->', ' wp:', '<!-- wp:', '<!-- /wp:' );
$seeds = array_merge( $corpus['gb-fixtures'], $corpus['small-post'] );
$seeds = array_values( array_filter( $seeds, function ( $d ) { return strlen( $d ) < 4000; } ) );
for ( $m = 0; $m < $mutations; $m++ ) {
$doc = $seeds[ mt_rand( 0, count( $seeds ) - 1 ) ];
$edits = mt_rand( 1, 4 );
for ( $e = 0; $e < $edits; $e++ ) {
$at = mt_rand( 0, strlen( $doc ) );
$op = mt_rand( 0, 2 );
$len = mt_rand( 1, 3 );
$ins = $alphabet[ mt_rand( 0, count( $alphabet ) - 1 ) ];
if ( 0 === $op ) {
$doc = substr( $doc, 0, $at ) . $ins . substr( $doc, $at );
} elseif ( 1 === $op ) {
$doc = substr( $doc, 0, $at ) . substr( $doc, $at + $len );
} else {
$doc = substr( $doc, 0, $at ) . $ins . substr( $doc, $at + $len );
}
}
if ( 0 === mt_rand( 0, 3 ) ) {
$doc = substr( $doc, 0, mt_rand( 0, strlen( $doc ) ) );
}
$docs[ "mut#$m" ] = $doc;
}
echo "\nDifferential trace, trunk vs PR\n";
$modes = array( 'token', 'delim', 'delim*', 'delim-ff', 'delim-p', 'block', 'block*', 'block-group', 'extract', 'extract-inner', 'mixed', 'delim-ns', 'delim-core-ns', 'delim-odd' );
$failures = 0;
$checked = 0;
foreach ( $docs as $name => $doc ) {
foreach ( $modes as $mode ) {
$a = trace( $processors['trunk'], $doc, $mode );
$b = trace( $processors['PR'], $doc, $mode );
++$checked;
if ( $a === $b ) {
continue;
}
if ( ++$failures <= 10 ) {
foreach ( $a as $k => $v ) {
if ( ! array_key_exists( $k, $b ) || $b[ $k ] !== $v ) {
echo " DIFF $name [$mode] step $k\n doc: " . json_encode( substr( $doc, 0, 300 ) ) . "\n trunk: " . json_encode( $v ) . "\n PR: " . json_encode( $b[ $k ] ?? '(missing)' ) . "\n";
break;
}
}
if ( count( $a ) !== count( $b ) && array_slice( $a, 0, count( $b ) ) === $b ) {
echo " DIFF $name [$mode] length " . count( $a ) . ' vs ' . count( $b ) . "\n";
}
}
}
}
printf( " %d documents, %d traces, %d differ\n", count( $docs ), $checked, $failures );
unset( $docs );
}
/*
* Benchmark.
*/
$works = array(
'token' => 'next_token()',
'block' => 'next_block()',
'type' => 'next_block( $type )',
'extract' => 'extract',
'parse' => 'parse',
);
function run_work( string $work, string $class, array $docs ): int {
$count = 0;
switch ( $work ) {
case 'token':
foreach ( $docs as $doc ) {
$p = new $class( $doc );
while ( $p->next_token() ) {
++$count;
}
}
break;
case 'block':
foreach ( $docs as $doc ) {
$p = new $class( $doc );
while ( $p->next_block() ) {
$p->get_block_type();
++$count;
}
}
break;
case 'type':
foreach ( $docs as $doc ) {
foreach ( array( 'core/heading', 'image', 'my-plugin/missing' ) as $type ) {
$p = new $class( $doc );
while ( $p->next_block( $type ) ) {
++$count;
}
}
}
break;
case 'extract':
foreach ( $docs as $doc ) {
$p = new $class( $doc );
while ( $p->next_block( '*' ) ) {
$p->extract_full_block_and_advance();
++$count;
}
}
break;
case 'parse':
foreach ( $docs as $doc ) {
$count += count( ( new $class() )->parse( $doc ) );
}
break;
}
return $count;
}
if ( null !== $child ) {
if ( 'parse' === $child_work ) {
$class = 'regex' === $child ? 'WP_Block_Parser_Regex' : "WP_Block_Parser_13695_$child";
} else {
$class = $processors[ $child ];
}
$results = array();
foreach ( $corpus as $set => $docs ) {
// Repeat each pass to fill about 40 ms: once to warm up the JIT, once to measure.
$t = hrtime( true );
$count = run_work( $child_work, $class, $docs );
$n = max( 1, (int) ceil( 0.04 / ( max( 1, hrtime( true ) - $t ) / 1e9 ) ) );
for ( $i = 0; $i < $n; $i++ ) {
run_work( $child_work, $class, $docs );
}
$t = hrtime( true );
for ( $i = 0; $i < $n; $i++ ) {
run_work( $child_work, $class, $docs );
}
$results[ $set ] = array( ( hrtime( true ) - $t ) / $n / 1e3, $count );
}
echo json_encode( $results );
exit( 0 );
}
/**
* Returns the PHP command for a child process, with this process's opcache and JIT settings.
*/
function child_command(): string {
$php = escapeshellarg( PHP_BINARY );
$fresh = json_decode( (string) shell_exec( "$php -r " . escapeshellarg( 'echo json_encode( extension_loaded( "Zend OPcache" ) ? ini_get_all( "zend opcache", false ) : null );' ) ), true );
$cmd = $php;
if ( ! is_array( $fresh ) ) {
$cmd .= ' -d zend_extension=opcache';
$fresh = array();
}
foreach ( ini_get_all( 'zend opcache', false ) ?: array() as $key => $value ) {
if ( ! array_key_exists( $key, $fresh ) || $fresh[ $key ] !== $value ) {
$cmd .= ' -d ' . escapeshellarg( "$key=$value" );
}
}
return $cmd;
}
if ( $run_bench ) {
$command = child_command() . ' ' . escapeshellarg( __FILE__ );
foreach ( array( $trunk, $pr, $gutenberg ) as $path ) {
$command .= ' ' . escapeshellarg( $path );
}
if ( null !== $parser ) {
$command .= ' ' . escapeshellarg( "--parser=$parser" );
}
$jobs = array();
foreach ( $works as $work => $label ) {
if ( 'parse' === $work ) {
if ( null !== $parser ) {
$jobs[] = array( $work, array( 'regex', 'trunk', 'PR' ) );
}
} else {
$jobs[] = array( $work, array( 'trunk', 'PR' ) );
}
}
// $times[ $work ][ $set ][ $tag ] is a list of µs per pass, one per round.
$times = array();
$counts = array();
fwrite( STDERR, "\nRounds: " );
for ( $r = 0; $r < $rounds; $r++ ) {
foreach ( $jobs as list( $work, $tags ) ) {
shuffle( $tags );
foreach ( $tags as $tag ) {
$cmd = $command . ' ' . escapeshellarg( "--child=$tag" ) . ' ' . escapeshellarg( "--work=$work" );
$out = json_decode( (string) shell_exec( $cmd ), true );
if ( ! is_array( $out ) ) {
fwrite( STDERR, "\nThe process for $tag $work failed: $cmd\n" );
exit( 2 );
}
foreach ( $out as $set => list( $us, $count ) ) {
$times[ $work ][ $set ][ $tag ][] = $us;
$counts[ $work ][ $set ][ $tag ] = $count;
}
}
}
fwrite( STDERR, ( $r + 1 ) . ' ' );
}
fwrite( STDERR, "\n" );
$median = function ( array $list ): float {
sort( $list );
return $list[ intdiv( count( $list ), 2 ) ];
};
foreach ( $counts as $work => $sets ) {
foreach ( $sets as $set => $by_tag ) {
if ( count( array_unique( $by_tag ) ) > 1 ) {
echo " MISMATCH $set {$works[ $work ]}: " . json_encode( $by_tag ) . "\n";
}
}
}
printf( "\nWP_Block_Processor, median of %d rounds, µs per pass over the set\n", $rounds );
printf( " %-16s %-20s %10s %10s %8s\n", 'set', 'work', 'trunk', 'PR', 'PR/trunk' );
foreach ( array_keys( $corpus ) as $set ) {
foreach ( $works as $work => $label ) {
if ( 'parse' === $work ) {
continue;
}
$a = $median( $times[ $work ][ $set ]['trunk'] );
$b = $median( $times[ $work ][ $set ]['PR'] );
printf( " %-16s %-20s %10.1f %10.1f %8.3f\n", $set, $label, $a, $b, $b / $a );
}
}
if ( null !== $parser ) {
printf( "\nWP_Block_Parser::parse(), median of %d rounds, µs per pass over the set\n", $rounds );
printf( " %-16s %10s %10s %10s %8s %8s %8s\n", 'set', 'regex', '#13695+tr', '#13695+PR', 'tr/regex', 'PR/regex', 'PR/tr' );
foreach ( array_keys( $corpus ) as $set ) {
$regex = $median( $times['parse'][ $set ]['regex'] );
$on_trunk = $median( $times['parse'][ $set ]['trunk'] );
$on_pr = $median( $times['parse'][ $set ]['PR'] );
printf( " %-16s %10.1f %10.1f %10.1f %8.3f %8.3f %8.3f\n", $set, $regex, $on_trunk, $on_pr, $on_trunk / $regex, $on_pr / $regex, $on_pr / $on_trunk );
}
}
}
exit( $run_diff && $failures ? 1 : 0 ); |
Test using WordPress PlaygroundThe changes in this pull request can previewed and tested using a WordPress Playground instance. WordPress Playground is an experimental project that creates a full WordPress instance entirely within the browser. Some things to be aware of
For more details about these limitations and more, check out the Limitations page in the WordPress Playground documentation. |
|
I measured each commit in isolation, as trunk plus that commit and as this branch without it. Commits 2, 6, 7 and 8 do not apply to trunk alone, and removing commit 1 or 2 from this branch changes the commits after it, so those variants are ported by hand. Every variant matches trunk in the differential trace (1,451,408 traces, 0 differ) and passes this branch's 398 block-processor tests. Removing each commit from this branch, PHP 8.5 with opcache, 20 rounds, each variant in its own process. Percent more time on block content without the commit, JIT off / tracing JIT; negative means faster without it:
Commit 9 rewrites commit 4's loop, so they are removed together. Without commit 3, every operation on the classic post takes 4.8 to 5 times as long; no other commit changes it. Commit 5 changes only docblocks.
Most 95% confidence intervals are within ±0.5% without JIT and ±1% with it. |
Possible
WP_Block_Processorperformance optimizations, relevant for #13695. This reduces the processor's cost per token and per delimiter. Every public method returns the same values as before.next_delimiter()passes over HTML spans insidenext_token()and normalizes a searched block type once instead of on every delimiter.!--, and finds a delimiter's comment closer in place by its>instead of callingfind_html_comment_end().extract_full_block_and_advance()read the processor's state instead of calling other public methods.@phpstan-impure.Each commit message gives its measured effect. Time on this branch divided by time on trunk, PHP 8.5 with opcache, no JIT unless noted:
WP_Block_Parser::parse()from #13695next_block()next_block( $type )extract_full_block_and_advance()Block content is Gutenberg's two performance posts, its block fixtures, and the block themes' templates and patterns. The classic post is the large performance post without block delimiters.
A differential trace compared 20 accessors after every step of 14 traversal modes over that content, edge cases, truncations and 100,000 mutated documents: 1,451,408 traces, no differences. A subclass that reimplements public methods sees fewer internal calls:
next_delimiter()no longer callsis_block_type()when searching for a type other than*or freeform, extraction no longer callsget_depth(),get_html_content()oropens_block(), andget_printable_block_type()now callsget_block_type(). The class documentation supports reimplementing onlyget_last_error(),get_attributes()andget_last_json_error(), and no plugin in the directory subclasses it.New tests cover breadcrumbs and depth on void blocks and HTML,
next_delimiter()against a filterednext_token(),!--outside comments, dash runs in block names,>in attributes and null block types. The first comment has a script that reproduces the trace and the timings.Trac ticket:
Use of AI Tools
AI assistance: Yes
Tool(s): Claude Code
Model(s): Claude Opus 5.5
Used for: Implementation, tests, benchmarks, the differential trace and this description; reviewed by me.
This Pull Request is for code review only. Please keep all other discussion in the Trac ticket. Do not merge this Pull Request. See GitHub Pull Requests for Code Review in the Core Handbook for more details.