Guide
How to fix URIError: URI malformed in JavaScript
Find incomplete percent escapes, invalid UTF-8 bytes, and lone Unicode surrogates that cause URI malformed errors.
by Tools in a Tab · Published on · Updated
Short answer
URIError: URI malformed usually has one of two causes. During decoding, a %
escape is incomplete or its bytes are not valid UTF-8. During encoding, the
JavaScript string contains a lone Unicode surrogate. Do not hide the exception
with an empty try/catch; preserve the input, identify the operation, and fix
the layer that produced it.
Fast diagnosis
| Operation | Typical failing input | Cause | Correct action |
|---|---|---|---|
decodeURIComponent or decodeURI |
%, %2, %GG |
Incomplete or non-hexadecimal escape | Fix or reject the producer’s escaped value |
decodeURIComponent or decodeURI |
%C3%28 |
Escapes are complete but bytes are not UTF-8 | Restore the correct bytes or character encoding |
encodeURIComponent or encodeURI |
Lone surrogate | The JavaScript string is not well-formed | Preserve character boundaries or reject it |
Decoding a complete URL as one component is a separate mistake: %26, %3F,
and %23 can become structural delimiters without throwing any exception.
Parse the address before reading individual values.
Start with the operation named in the stack trace. Changing from encoding to
decoding, replacing every %, or repeatedly decoding does not repair the
original boundary error.
Separate encoding from decoding first
decodeURIComponent and decodeURI consume %HH sequences. They fail on %,
%A, %GG, or byte sequences that do not form valid UTF-8. For example,
%E0%A4%A is truncated before the character can be completed.
encodeURIComponent and encodeURI consume Unicode text. They may fail when
the string contains half of a surrogate pair, which can happen after slicing a
UTF-16 string at an arbitrary index or receiving corrupted internal data.
The URL encoder and decoder separates component, complete URL, and form-value modes and reports errors without sending the input elsewhere.
Incomplete percent escapes
Every % must be followed by exactly two hexadecimal digits. %20 is a space
and %2F represents /; %2, %XZ, and a final % are invalid. Inspect the
original before replacing % with %25, because that operation can turn
broken input into the literal text of a broken escape rather than restore the
intended value.
A common example interpolates a human percentage such as 50% into a URL and
later tries to decode it. When it belongs to a value, encode it as 50%25 at
the boundary where that component is constructed.
You can detect malformed escape syntax before decoding while still allowing the decoder to reject invalid UTF-8:
function decodeComponentStrict(input) {
const malformedEscape = input.match(/%(?![0-9A-Fa-f]{2})/);
if (malformedEscape) {
throw new URIError(
`Incomplete percent escape at index ${malformedEscape.index}`,
);
}
return decodeURIComponent(input);
}
This does not “sanitize” a bad value. It separates an incomplete %HH escape
from a later UTF-8 failure and keeps the rejection visible to the caller.
Invalid UTF-8 byte sequences
Escapes can contain valid hexadecimal pairs and still be invalid together.
decodeURIComponent('%C3%28') fails because C3 begins a two-byte UTF-8
sequence but 28 is not a continuation byte. Do not map each %HH directly to
a Latin-1 character. Assemble the bytes and decode them as strict UTF-8.
MDN describes both forms in its URI malformed error reference. The actual repair often belongs in the producer that truncated the data or used a different character encoding.
Lone Unicode surrogates during encoding
JavaScript strings contain UTF-16 code units. Many characters outside the
basic plane, including emoji, require a high and low surrogate. Cutting between
them leaves a lone unit that represents no valid Unicode scalar value, so
encodeURI may throw URIError.
Avoid truncating text by raw code units when character boundaries matter.
Array.from(text) or for…of iterates by code point, although complex visible
symbols can still require grapheme segmentation. On supported runtimes,
text.isWellFormed() can detect lone surrogates explicitly.
When isWellFormed() is available, validate before encoding:
if (!input.isWellFormed()) {
throw new URIError('The input contains a lone Unicode surrogate.');
}
const encoded = encodeURIComponent(input);
Replacing a lone surrogate with � changes the data. That may be an explicit
display policy, but it is not a lossless repair.
A debugging procedure
- Record safely whether the failure happened during encoding or decoding, without logging sensitive values.
- For decoding, locate the first
%not followed by two hexadecimal digits. - If escape syntax is complete, assemble the bytes and validate UTF-8.
- For encoding, verify that the Unicode string is well formed.
- Confirm whether the function should handle one component or a complete URL.
Do not repeatedly call decodeURIComponent until the string stops changing.
Every layer needs a documented reason, and a second pass can turn encoded data
into active delimiters.
decodeURI or decodeURIComponent?
Use decodeURIComponent for one encoded value or path segment. Use decodeURI
only for a complete URI whose structural delimiters must remain delimiters.
Neither function fetches the address or proves that its scheme, host, or path
is allowed. When query parameters are involved, parse with URL and read them
through URLSearchParams instead of decoding the entire address first.
URLSearchParams already decodes query values using form rules, including
+ as a space. Do not decode the returned value again by default. It is also
forgiving of malformed escapes and UTF-8, so successful parsing is not strict
input validation. If rejection is required, validate the raw encoded value
first. The URL Standard
defines these forgiving form-parsing rules, which differ from strict URI decoding.
Prevention
Keep values unencoded inside the application and encode them once at their
insertion boundary. Build addresses with URL and URLSearchParams, validate
external input before decoding, and test accents, emoji, %, truncated
escapes, and invalid byte sequences. A visible error is safer than continuing
with a URL whose meaning has silently changed.
Frequently asked questions
Why does %E0%A4%A produce URI malformed?
The last escape is incomplete: %A lacks a second hexadecimal digit. Even
when every escape has two digits, the combined bytes must form valid UTF-8.
Should I replace every percent sign with %25?
No. That changes an escape into literal percent text and can hide corruption. Encode a raw component once at its insertion boundary; reject or repair the producer when an already encoded value is malformed.
Is try/catch the fix?
It is part of error handling, not data repair. Catch the exception to show a useful message or reject the request, but retain the reason and correct the layer that created the malformed value.