URL Parser

Split a URL into its component parts.

Input
URL input
Output
Result
Options

Spaces per level of nesting.

About this tool


A long URL with encoded parameters is hard to read and easy to misjudge. This parser breaks it into labelled components (protocol, host, port, path segments, each query parameter, and the fragment) using the browser's own URL implementation, so the result matches exactly how a browser would interpret it.

That last point is why parsing with a regular expression is a mistake. URLs have genuinely intricate rules around credentials, default ports, IPv6 hosts and relative resolution, and the built-in parser already implements the WHATWG specification correctly.

How to use it

  1. Paste or upload your urlDrop a file onto the input pane, use the file picker, or paste the text directly.
  2. Adjust the options if neededThe defaults suit most input; open Options to change the behaviour.
  3. ParsePress Parse, or use Ctrl+Enter (Cmd+Enter on macOS).
  4. Copy or downloadCopy the result, or download it as a .json file.

Worked examples


Each example below is executed against this tool by the test suite, so what you see is what the tool actually produces.

A URL with a port, query and fragment

Input

https://api.example.com:8443/v1/users?page=2&sort=name#results

Output

{
  "href": "https://api.example.com:8443/v1/users?page=2&sort=name#results",
  "protocol": "https",
  "username": null,
  "password": null,
  "hostname": "api.example.com",
  "port": "8443",
  "origin": "https://api.example.com:8443",
  "pathname": "/v1/users",
  "pathSegments": [
    "v1",
    "users"
  ],
  "search": "?page=2&sort=name",
  "queryParameters": {
    "page": "2",
    "sort": "name"
  },
  "hash": "results"
}

The path is also split into segments, and each query parameter is decoded separately.

What to watch for


The details that decide whether a conversion is correct, and where information can be lost without any error being raised.

Query parameters are decoded and grouped
Each parameter is percent-decoded and listed separately, so you can read values that were unreadable in the raw URL. A parameter repeated more than once becomes an array, which is the correct interpretation, a URL may legitimately carry the same key several times, and most frameworks expose those as a list.
Passwords are redacted
A URL can embed credentials as user:password@host. Since this output may be pasted into a ticket or a chat, any password is replaced with a placeholder in every field, including the full href, echoing it back would defeat the point of redacting it at all.
Default ports disappear
The port field is empty for https on 443 or http on 80, because the URL specification drops the default. That is not a parsing failure; it reflects that the URL is equivalent with or without it.
The fragment never reaches the server
Everything after # is handled by the browser alone and is not included in the HTTP request. This matters when debugging: a value in the fragment is invisible to server logs, which is why OAuth implicit flows that returned tokens there were so hard to audit.
A scheme is required
Parsing needs an absolute URL, so example.com/path alone is rejected, without a scheme, the string is a relative reference whose meaning depends on a base URL. Prefix it with https:// to parse.

Limitations


  • Requires an absolute URL including the scheme.
  • Passwords are redacted rather than displayed.
  • Processing happens in your browser, so very large inputs are bounded by available memory. Files above roughly 10 MB are handled but will feel slower, and multi-hundred-megabyte files are better suited to a command-line tool.

Questions


Why is the port empty?
Because it is the default for the scheme, 443 for https, 80 for http. The URL specification omits default ports, so the URL means the same thing either way.
Why is a repeated parameter shown as an array?
Because URLs permit the same key more than once, and dropping duplicates would misrepresent the URL. Most server frameworks expose these as a list too.
Why must I include https://?
A URL without a scheme is a relative reference, and its meaning depends on the page it appears on. Parsing requires an absolute URL.