Guide

How to fix URIError: URI malformed in JavaScript

Find incomplete percent escapes, invalid UTF-8 bytes, and lone Unicode surrogates that cause URI malformed errors.

by Tools in a Tab · Published on · Updated

Short answer

URIError: URI malformed usually has one of two causes. During decoding, a % escape is incomplete or its bytes are not valid UTF-8. During encoding, the JavaScript string contains a lone Unicode surrogate. Do not hide the exception with an empty try/catch; preserve the input, identify the operation, and fix the layer that produced it.

Fast diagnosis

Operation Typical failing input Cause Correct action
decodeURIComponent or decodeURI %, %2, %GG Incomplete or non-hexadecimal escape Fix or reject the producer’s escaped value
decodeURIComponent or decodeURI %C3%28 Escapes are complete but bytes are not UTF-8 Restore the correct bytes or character encoding
encodeURIComponent or encodeURI Lone surrogate The JavaScript string is not well-formed Preserve character boundaries or reject it

Decoding a complete URL as one component is a separate mistake: %26, %3F, and %23 can become structural delimiters without throwing any exception. Parse the address before reading individual values.

Start with the operation named in the stack trace. Changing from encoding to decoding, replacing every %, or repeatedly decoding does not repair the original boundary error.

Separate encoding from decoding first

decodeURIComponent and decodeURI consume %HH sequences. They fail on %, %A, %GG, or byte sequences that do not form valid UTF-8. For example, %E0%A4%A is truncated before the character can be completed.

encodeURIComponent and encodeURI consume Unicode text. They may fail when the string contains half of a surrogate pair, which can happen after slicing a UTF-16 string at an arbitrary index or receiving corrupted internal data.

The URL encoder and decoder separates component, complete URL, and form-value modes and reports errors without sending the input elsewhere.

Incomplete percent escapes

Every % must be followed by exactly two hexadecimal digits. %20 is a space and %2F represents /; %2, %XZ, and a final % are invalid. Inspect the original before replacing % with %25, because that operation can turn broken input into the literal text of a broken escape rather than restore the intended value.

A common example interpolates a human percentage such as 50% into a URL and later tries to decode it. When it belongs to a value, encode it as 50%25 at the boundary where that component is constructed.

You can detect malformed escape syntax before decoding while still allowing the decoder to reject invalid UTF-8:

function decodeComponentStrict(input) {
  const malformedEscape = input.match(/%(?![0-9A-Fa-f]{2})/);
  if (malformedEscape) {
    throw new URIError(
      `Incomplete percent escape at index ${malformedEscape.index}`,
    );
  }

  return decodeURIComponent(input);
}

This does not “sanitize” a bad value. It separates an incomplete %HH escape from a later UTF-8 failure and keeps the rejection visible to the caller.

Invalid UTF-8 byte sequences

Escapes can contain valid hexadecimal pairs and still be invalid together. decodeURIComponent('%C3%28') fails because C3 begins a two-byte UTF-8 sequence but 28 is not a continuation byte. Do not map each %HH directly to a Latin-1 character. Assemble the bytes and decode them as strict UTF-8.

MDN describes both forms in its URI malformed error reference. The actual repair often belongs in the producer that truncated the data or used a different character encoding.

Lone Unicode surrogates during encoding

JavaScript strings contain UTF-16 code units. Many characters outside the basic plane, including emoji, require a high and low surrogate. Cutting between them leaves a lone unit that represents no valid Unicode scalar value, so encodeURI may throw URIError.

Avoid truncating text by raw code units when character boundaries matter. Array.from(text) or for…of iterates by code point, although complex visible symbols can still require grapheme segmentation. On supported runtimes, text.isWellFormed() can detect lone surrogates explicitly.

When isWellFormed() is available, validate before encoding:

if (!input.isWellFormed()) {
  throw new URIError('The input contains a lone Unicode surrogate.');
}

const encoded = encodeURIComponent(input);

Replacing a lone surrogate with � changes the data. That may be an explicit display policy, but it is not a lossless repair.

A debugging procedure

  1. Record safely whether the failure happened during encoding or decoding, without logging sensitive values.
  2. For decoding, locate the first % not followed by two hexadecimal digits.
  3. If escape syntax is complete, assemble the bytes and validate UTF-8.
  4. For encoding, verify that the Unicode string is well formed.
  5. Confirm whether the function should handle one component or a complete URL.

Do not repeatedly call decodeURIComponent until the string stops changing. Every layer needs a documented reason, and a second pass can turn encoded data into active delimiters.

decodeURI or decodeURIComponent?

Use decodeURIComponent for one encoded value or path segment. Use decodeURI only for a complete URI whose structural delimiters must remain delimiters. Neither function fetches the address or proves that its scheme, host, or path is allowed. When query parameters are involved, parse with URL and read them through URLSearchParams instead of decoding the entire address first.

URLSearchParams already decodes query values using form rules, including + as a space. Do not decode the returned value again by default. It is also forgiving of malformed escapes and UTF-8, so successful parsing is not strict input validation. If rejection is required, validate the raw encoded value first. The URL Standard defines these forgiving form-parsing rules, which differ from strict URI decoding.

Prevention

Keep values unencoded inside the application and encode them once at their insertion boundary. Build addresses with URL and URLSearchParams, validate external input before decoding, and test accents, emoji, %, truncated escapes, and invalid byte sequences. A visible error is safer than continuing with a URL whose meaning has silently changed.

Frequently asked questions

Why does %E0%A4%A produce URI malformed?

The last escape is incomplete: %A lacks a second hexadecimal digit. Even when every escape has two digits, the combined bytes must form valid UTF-8.

Should I replace every percent sign with %25?

No. That changes an escape into literal percent text and can hide corruption. Encode a raw component once at its insertion boundary; reject or repair the producer when an already encoded value is malformed.

Is try/catch the fix?

It is part of error handling, not data repair. Catch the exception to show a useful message or reject the request, but retain the reason and correct the layer that created the malformed value.