Latest Results
fix(linter): search string values without UTF-8 conversion
Several rules converted a string value to UTF-8 before looking for a
character or prefix, so a lone surrogate anywhere in the value hid
text elsewhere in it. jest/valid-title missed the surrounding spaces
in " \uD800 " and the duplicate prefix in "test \uD800", and
no-template-curly-in-string, no-script-url, no-regex-spaces,
no-webpack-loader-syntax, node/no-sync, and
jest/prefer-lowercase-title missed the text they check next to a
surrogate.
Add contains, starts_with, ends_with, find, and rfind on JSStr with
the argument forms consumers use, &str, char, and char predicates,
through a sealed JSStrPattern trait. When the value has no lone
surrogate the methods delegate to str; otherwise text patterns are
searched byte for byte in canonical WTF-8, which finds exactly the
UTF-16 matches because a UTF-8 pattern cannot start or end inside
another code point's encoding and never equals a lone surrogate's
encoding. Predicates decode code points and never match a lone
surrogate, since no char exists for it; the trait documents that a
negated predicate therefore does not see every code point.
Use these methods in the rules above and in the formatter's JSX
attribute multiline checks. Where a rule still needs UTF-8, the
conversion is confined to that step: prefer-lowercase-title reports
without its UTF-8 fix, and no-sync and no-webpack-loader-syntax render
the name with escapes in the message. A title that begins with a lone
surrogate has no case, so prefer-lowercase-title accepts it.
Tests compare every method and pattern form with str on UTF-8 input
and with a UTF-16 reference model around lone halves and pairs formed
across a boundary, and cover each rule with a surrogate beside the
matched text and as the only content.
Part of #26242.
Implemented with AI assistance. feat(ecmascript): fold indexOf, lastIndexOf, and charCodeAt on strings with lone surrogates
A string literal containing a lone surrogate has no UTF-8
representation, so every string method fold declined it and calls like
"a\uD800b".indexOf("b") or "a\uD800b".charCodeAt(1) survived
minification. The numeric folds do not need a UTF-8 receiver: indexOf,
lastIndexOf, and charCodeAt return positions and code units, which are
defined for any sequence of UTF-16 code units.
Implement StringIndexOf, StringLastIndexOf, and StringCharCodeAt for
JSStr. A value without lone surrogates delegates to the existing UTF-8
implementations. Otherwise the search compares UTF-16 code units
directly, the domain the specification defines, so a lone surrogate
counts as one unit and no boundary rounding is needed. JSStr keeps its
WTF-8 bytes private, which rules out the byte search the UTF-8 path
uses; the unit walk iterates the encoded units without allocating and
only runs on the rare lone-surrogate literals. A search value with a
lone surrogate still declines, since constant string values remain
UTF-8. charAt also still declines: its result is a string, which may
have no UTF-8 form.
Verified against node on a generated corpus of 4382 string method
calls over lone and paired surrogate values, with identical program
output, byte-identical double minification, and unchanged minifier
conformance snapshots.
Part of #26242.
Implemented with AI assistance.jsstr-surrogate-numeric-folds Latest Branches
+24%
0%
0%
Ā© 2026 CodSpeed Technology