Skip to content

fix(url_decode): stop malformed escapes failing the render - #970

Open
josefguenther wants to merge 2 commits into
harttle:masterfrom
josefguenther:fix/url-decode-malformed-escape
Open

josefguenther wants to merge 2 commits into
harttle:masterfrom
josefguenther:fix/url-decode-malformed-escape

Conversation

@josefguenther

@josefguenther josefguenther commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

url_decode (src/filters/url.ts:3) passes its whole input to decodeURIComponent, so a single malformed escape throws and fails the render. {{ v | url_decode }}:

v master Shopify render this PR
100% RenderError 100% 100%
%zz RenderError %zz %zz
%E2%9C%93 done 100% RenderError ✓ done 100% ✓ done 100%
a%2Bb 100% RenderError a+b 100% a+b 100%
%E2%9C RenderError Liquid error: invalid byte sequence in UTF-8 �
%ff RenderError Liquid error: invalid byte sequence in UTF-8 �
caf%C3%A9%ff RenderError Liquid error: invalid byte sequence in UTF-8 café�
caf%C3%A9, a+b%20c, 1%2B1 unchanged unchanged

This decodes each run of %XX escapes as UTF-8 bytes, so:

  • a % that isn't followed by two hex digits stays as written, as in Shopify;
  • a byte sequence that isn't valid UTF-8 becomes U+FFFD, as the URL Standard's percent-decoding (and URLSearchParams) does.

Shopify's url_decode raises on invalid UTF-8 instead. Its default render writes that error inline and carries on, while LiquidJS fails the whole render, so this takes the URL Standard's answer.

For input that decodes today the output is unchanged. + still becomes a space before decoding, so #939's %2B handling stays. The docs page for the filter now states how malformed input renders.

Four new tests in test/integration/filters/url.spec.ts: a stray %, an incomplete sequence, and valid and invalid bytes in one run fail on master and pass here; a leading byte order mark passes on both and guards against the decoder dropping it. npm run build, npm test (1676 tests), npm run lint and npm run build:docs pass on Node 24.11.1.

Fixes #966

Written with AI assistance (Claude Code); I have reviewed the change.

url_decode passed its whole input to decodeURIComponent, so a single
malformed escape (a "%" not followed by two hex digits, or an
incomplete UTF-8 sequence) threw a RenderError and failed the render,
discarding any valid escapes around it. It now decodes each run of
well-formed escapes separately and leaves the rest as written, which
is how Shopify's CGI.unescape treats a stray "%". "+" is still turned
into a space before decoding, so %2B stays a literal plus.

Fixes harttle#966

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@harttle harttle added the bug label Oct 1, 2026
@harttle

harttle commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

I find one discrepency:

input                           ruby              liquidjs          PR #970
"100%"                          "100%"            error             "100%"
"%zz"                           "%zz"             error             "%zz"
"%E2%9C"                        error             error             "%E2%9C"
"%E2%9C%93 done 100%"           "✓ done 100%"     error             "✓ done 100%"
"a%2Bb 100%"                    "a+b 100%"        error             "a+b 100%"
"caf%C3%A9"                     "café"            "café"            "café"
"a+b c"                         "a b c"           "a b c"           "a b c"
"1%2B1"                         "1+1"             "1+1"             "1+1"
"foo+bar"                       "foo bar"         "foo bar"         "foo bar"
"foo%20bar"                     "foo bar"         "foo bar"         "foo bar"
"foo%2B1@example.com"           "foo+1@example.com"  "foo+1@example.com"  "foo+1@example.com"
"%E2%9C%93+100%"                "✓ 100%"          error             "✓ 100%"
"%ff"                           error             error             "%ff"

Not sure for other cases, but you can check with Shopify/liquid.

Each run of escapes was decoded with decodeURIComponent and kept as
written if that threw, so one byte that isn't valid UTF-8 left its
whole run undecoded: "caf%C3%A9%ff" rendered as written rather than
"café�". Each run is now decoded as UTF-8 bytes, and a byte sequence
that forms no character becomes U+FFFD, as the URL Standard's
percent-decoding does. A leading byte order mark is kept, as
decodeURIComponent keeps it.

The docs page now states how malformed input renders.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@josefguenther josefguenther changed the title fix(url_decode): leave malformed escapes as written fix(url_decode): stop malformed escapes failing the render Oct 1, 2026
@josefguenther

Copy link
Copy Markdown
Contributor Author

Thanks for checking against Ruby. You're right about those two rows, and they led me to a worse case the first commit got wrong: one bad byte left its whole run of escapes undecoded, so caf%C3%A9%ff came out as caf%C3%A9%ff rather than café�. The new commit decodes each run as UTF-8 bytes and replaces only the bytes that don't form a character with U+FFFD, the way the URL Standard and URLSearchParams do. A stray % still stays as written, like Shopify.

The invalid-UTF-8 rows are the same question as base64 in #971; I've written up there why I think the render should continue.

Comment thread src/filters/url.ts
export const url_decode = (x: string) => stringify(x)
.replace(/\+/g, ' ')
.replace(rPercentEscapes, escapes => new TextDecoder('utf-8', { ignoreBOM: true })
.decode(Uint8Array.from(escapes.slice(1).split('%'), hex => parseInt(hex, 16))))

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for the Uint8Array and TextDecoder, need add test on both Node and browser bundles.

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They're not as old and stable as decodeURIComponent

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

url_decode throws on a malformed % escape

2 participants