replace raw null and surrogate code points during tokenization - #191
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #191 +/- ##
=======================================
Coverage 99.79% 99.79%
=======================================
Files 3 3
Lines 966 968 +2
Branches 155 155
=======================================
+ Hits 964 966 +2
Misses 1 1
Partials 1 1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
|
makes sense on both. trimmed the comments down to a single §3.3 line (regex + call site) and cut the lxml bit from the test comment. reworded the description too, since you're right that §3.3 only covers u+0000 and surrogates, so the raw control-char case is out of scope and this is really a conformance fix, not an lxml-compat one. |
tokenize() consumes the selector string without the input-stream preprocessing step from css syntax level 3 §3.3, so a raw u+0000 or lone surrogate code point that appears directly in a string token (not as an escape) survives through attribute and :contains values into the final xpath. #164 and #189 only folded the escaped forms inside _replace_unicode (§4.3.7); the raw input path was never preprocessed.
this folds the raw u+0000 and surrogate forms to u+fffd at the entry of tokenize() with a length-preserving substitution so token positions stay correct, mirroring what the escape decoder already does. §3.3 only covers those code points, so this is a spec-conformance fix rather than a general lxml-compat one; other raw control characters are out of scope and still pass through.