Extracting Query Parameters from a URL
Query strings carry data in a URL's searchable component. Every character after the ? and before the # is part of the query string, and most applications rely on individual parameters embedded within it, including page numbers, search terms, filter selections, tracking tokens, and authentication state. Parsing them correctly means handling percent-encoded values, repeated keys, keys with empty values, and the difference between the raw string and its decoded form.
Every language provides at least one standard way to parse a query string without manual splitting. In JavaScript, URLSearchParams handles it.1 In Python, urllib.parse.parse_qs returns a dict of lists. In Go, url.Values carries all parameters. Using these built-in parsers avoids subtle bugs around percent-decoding and delimiter handling that trip up manual string splits.
Run this check yourself in the URL Parser & Inspector.
Open in the tool →How query strings are structured
After the ?, parameters follow the format key=value. Multiple pairs are separated by &. The same key can appear multiple times: ?tag=js&tag=python is a valid query string with two tag values. Consequently, any parser that stores results in a plain dict, where each key maps to one string, silently drops duplicate values. For APIs that use repeated keys as arrays, always use a parser that returns a list per key, like parse_qs() in Python or URLSearchParams.getAll() in JavaScript.1 Keys without a value, such as ?flag or ?key=, are also valid and may have distinct meanings. A bare key like ?flag is typically treated as a boolean presence flag, while ?key= represents a key with an empty string value.
Reading and modifying parameters
Reading a single parameter by name is a one-liner in every major language. JavaScript: url.searchParams.get('page'). Python: parse_qs(query)['page'][0]. PHP: parse_str($query, $out); echo $out['page']. Building on this, modifications follow the same API: in JavaScript, url.searchParams.set('page', '2') updates the parameter and url.searchParams.delete('tag') removes it, and the url.href reflects the change immediately.2 For languages without a mutable params object, rebuild the query string from a modified dict using the appropriate encoding function. In Python, calling urlencode() on the updated dict produces a new query string that you can assign to the URL. In Go, modifying the url.Values map and calling its Encode() method returns the updated query string in the correct format.
Percent-decoding, plus signs and bare keys
Percent-encoded values must be decoded before display or comparison: %20 is a space, %2B is a literal +, and + itself means a space in application/x-www-form-urlencoded format but is a literal + in RFC 3986 path contexts.3 Built-in parsers handle this automatically in most runtimes, so manual decoding is only needed when parsing query strings with a custom parser.
Handling keys with empty values versus bare keys
A key with an empty value (?key=) and a key with no value (?key) may both appear in the same query string, and different parsers treat them differently. Some parsers omit bare keys entirely, while others represent them with an empty string or a null value. Test your parser against both cases to confirm which behavior you get, since treating a bare key as a boolean flag while the parser returns an empty string leads to subtle logic bugs.
Parsing query strings in URL fragments for SPA routing
Single-page applications frequently encode route state in the fragment after a secondary hash or in a query string within the fragment. A URL like https://app.example.com/#/dashboard?tab=analytics&date=2026-06 contains the query parameters tab and date inside the fragment, not in the server-visible query string. JavaScript on the page reads window.location.hash to get #/dashboard?tab=analytics&date=2026-06, then parses that string as a query string using new URLSearchParams(hash.split('?')[1]). This pattern lets SPAs maintain navigation state without triggering a full page reload.
Working with nested query parameters across frameworks
Different server frameworks handle nested query parameters with bracket notation differently.4 PHP automatically parses bracket notation into nested arrays so that $_GET['filter']['status'] equals 'active'. Ruby on Rails converts filter[status] into params[:filter][:status] through its routing layer. Java's Servlet API does not parse brackets automatically, requiring a library or manual parsing. Python's parse_qs returns flat keys containing the literal bracket notation unless you post-process them into a nested structure yourself. When building an API consumed by multiple languages, document whether you support bracket notation or prefer dot-notation, and test parsing on each consumer.
Documenting the convention up front saves debugging time because the symptom of a mismatch is usually a silently empty nested field rather than an error. If you control both the client and server, pick one scheme and encode it consistently so the bracket or dot structure survives the round trip intact. When you must accept input from a third party, normalize their format into your internal representation before any business logic touches the parameters.
Query string security: parameter pollution and injection
Watching for HTTP parameter pollution in query strings matters because an application that receives multiple values for the same key but uses only the first or last without documenting which5 lets an attacker who appends ?role=admin to a URL where ?role=user already exists escalate privileges if the application reads the last value. Different frameworks choose differently: PHP uses the last value, Node.js's querystring.parse returns the last value, but URLSearchParams.get also returns the first.6 Building on this, always document which duplicate-value semantics your API uses and validate accordingly.
Preventing query parameter injection in server-side redirects
A server-side redirect that forwards query parameters from the request URL to a redirect target is vulnerable to parameter injection. If your /login endpoint redirects to /dashboard?returnTo=/home&theme=dark and the attacker crafts /login?returnTo=/admin&theme=<script>, the unencoded values flow into the redirect URL. Encoding all user-supplied values with encodeURIComponent() before concatenating them into a redirect URL prevents most injection attacks, but a stronger approach is to use an allowlist of permitted returnTo paths and ignore the user-supplied value entirely if it does not match. CapyToolkit runs all URL parsing examples locally in your browser for safe testing.
When to use this
Use a dedicated query parameter parser whenever your code reads parameters from user-supplied URLs, API responses, OAuth redirects, or webhook payloads, any time the query string may contain encoded characters, repeated keys, or empty values your code did not generate.
Examples
Parse a query string with repeated keys in JavaScript
// Manual split — loses duplicate tag values const query = "q=search&page=2&tag=js&tag=python"; const params = Object.fromEntries(new URLSearchParams(query)); console.log(params.tag); // "python" — "js" was lost
// Correct: use getAll() for repeated keys
const params2 = new URLSearchParams(query);
console.log(params2.getAll("tag")); // ["js", "python"] Parse query string with percent-encoded values in Python
from urllib.parse import parse_qs params = parse_qs("q=hello%20world&page=2&tag=python&tag=url") # params["q"] → ["hello world"] (decoded automatically) # params["tag"] → ["python", "url"] (both values preserved)
- 1.
Mozilla Developer Network, "URLSearchParams," developer.mozilla.org, accessed June 2026. https://developer.mozilla.org/en-US/docs/Web/API/URLSearchParams
- 2.
WHATWG, "URL Standard," url.spec.whatwg.org, accessed June 2026. https://url.spec.whatwg.org/
- 3.
"Query string," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/Query_string
- 4.
PHP, "parse_str," php.net, accessed June 2026. https://www.php.net/manual/en/function.parse-str.php
- 5.
OWASP, "Testing for HTTP Parameter Pollution," owasp.org, accessed June 2026. https://owasp.org/www-project-web-security-testing-guide/stable/4-Web_Application_Security_Testing/07-Input_Validation_Testing/04-Testing_for_HTTP_Parameter_Pollution
- 6.
Node.js, "Query string," nodejs.org, accessed June 2026. https://nodejs.org/docs/latest-v26.x/api/querystring.html
Extracting UTM Parameters from URLs
UTM parameters are query string keys that analytics platforms use to attribute web traffic to its source. The five standard parameters, utm_source, utm_medium, utm_campaign, utm_term, and utm_content, were popularized by Google Analytics and are now supported by every major analytics tool.1 Parsing them from incoming URLs is a common task in landing page handlers, analytics pipelines, and A/B testing frameworks.
Because UTM parameters are plain query string entries, they are parsed by the same APIs as any other query parameter. The key challenges are forwarding them across page navigations without losing them, persisting them in session storage for multi-page attribution, and handling URL-encoded values that contain spaces or special characters.
Run this check yourself in the URL Parser & Inspector.
Open in the tool →Reading UTM parameters
Parse the current page's UTM parameters in JavaScript with const url = new URL(window.location.href); const source = url.searchParams.get('utm_source'). This returns null if the parameter is absent, making it safe to call unconditionally. Building on this, collect all five parameters at once with a helper: const utm = Object.fromEntries(['utm_source','utm_medium','utm_campaign','utm_term','utm_content'].map(k => [k, url.searchParams.get(k)]).filter(([,v]) => v !== null)). On the server in Python, parse_qs(query)['utm_source'][0] extracts the first value; check for the key's presence before accessing it.2 In Go, iterate over url.Query() to read each UTM parameter, filtering for the utm_ prefix keys and ignoring any that are empty or missing. In PHP, the $_GET superglobal gives direct access to query parameters, so isset($_GET['utm_source']) checks for presence before reading the value. Parsing UTM parameters correctly means accounting for all five standard keys, handling missing values gracefully, and normalizing case sensitivity before the data reaches your analytics platform.
Persisting and forwarding UTM parameters
For single-page applications, store UTM parameters in sessionStorage on first load and read them when a conversion event fires, ensuring the attribution survives client-side navigations that replace the URL. Consequently, for multi-page sites, append the UTM parameters to all internal links using url.searchParams.set('utm_source', stored_source) before setting anchor hrefs. Server-side, persist UTM parameters in the user's session on the first page request and associate them with any conversion events during that session. Without this persistence, the attribution data is lost when the user navigates to a new page on your site. When forwarding UTM parameters to external services or analytics platforms, always validate that the values are properly encoded and do not contain characters that could break the receiving system.
Spaces, letter case and double-encoding in UTM values
UTM parameter values commonly contain spaces (encoded as %20 or +), hyphens, underscores, and slashes. Built-in URL parsers decode them automatically, so url.searchParams.get('utm_campaign') returns 'spring sale' for both %20 and + encoded values.2 Yet some analytics platforms are case-sensitive about UTM values: "Google" and "google" may be tracked as separate sources. Normalize values to lowercase on read if your analytics platform requires it. Building on this, watch for double-encoding: if your server reads a UTM value from a URL and then appends it to a redirect URL, encode it with encodeURIComponent() to prevent the spaces from breaking the query string.
Server-side UTM extraction in analytics pipelines
Analytics pipelines that process server logs or webhook events need to extract UTM parameters from raw URL strings. In a Node.js log processor, parse each request URL with new URL(logEntry.url) and read the searchParams. For Python-based ETL jobs, urlparse(log_url).query followed by parse_qs() gives you a dict of all query parameters including UTM keys. The key difference from browser-side extraction is that server-side code sees the raw URL before any client-side JavaScript has modified it, which means you capture the original attribution data even if the SPA later changes the URL.
Handling UTM parameters in webhook payloads
Many analytics platforms (Segment, Mixpanel, Amplitude) accept UTM parameters as part of their tracking payloads. When forwarding UTM data to these platforms, map the URL parameters to the expected field names: utm_source becomes traffic_type or referrer_source depending on the platform. Some platforms auto-extract UTM parameters from the page URL on the client side; others require you to pass them explicitly in the event properties. Check your platform's documentation to avoid double-counting the same attribution data from both automatic extraction and manual forwarding.
Keep the original URL parameter names alongside the platform-specific mapping so you can reconstruct the campaign if a vendor changes its schema. Store the raw UTM values in your own warehouse before they reach the third-party platform, because the platform's normalization may collapse variations you later want to analyze. This local copy also lets you reconcile discrepancies when two platforms report different attribution for the same session.
UTM parameters and privacy regulations
UTM parameters are not personally identifiable information on their own, but they become tracking data when combined with IP addresses, user agents, or session identifiers. Under GDPR and CCPA, the combination of UTM parameters with other session data may constitute tracking that requires user consent.3 Google Analytics 4 strips UTM parameters from the URL after processing them, storing only the derived session attribution.4 If your application stores UTM parameters in a database linked to user accounts, include them in your data processing inventory and privacy policy disclosures.
First-party UTM parameters versus third-party tracking
UTM parameters are a first-party tracking mechanism: you control the parameter values when you create the campaign links, and the data stays within your analytics platform. This contrasts with third-party tracking pixels and cross-site cookies, which are increasingly blocked by browser privacy features. UTM-based attribution continues to work in Safari's Intelligent Tracking Prevention and Firefox's Enhanced Tracking Protection because it relies on first-party URL parameters rather than third-party cookies.5 For privacy-conscious analytics stacks, UTM parameters combined with server-side session attribution provide a compliant alternative to third-party tracking.
Building UTM-tagged URLs at scale
Marketing teams that generate hundreds of campaign URLs need a systematic approach to UTM parameter management. A URL builder tool (spreadsheet, web form, or API) enforces consistent naming conventions: always lowercase utm_source values, use hyphens instead of spaces in utm_campaign, and maintain a controlled vocabulary for utm_medium (cpc, email, social, organic, referral). Without naming discipline, your analytics reports fragment across dozens of variations: "Email", "email", "e-mail", and "EMAIL" appear as four separate channels.
Validating UTM parameters before storing them
Validate UTM parameter values before writing them to your analytics database: flag an oversized utm_source as a likely injection attempt, normalize known source names to a canonical form, and strip whitespace from both ends. A validation function that checks utm_source against an allowlist of known traffic sources catches typos and prevents garbage data from polluting your reports. For utm_campaign, enforce a naming convention with a regular expression like /^[a-z0-9-]+$/ to prevent special characters from breaking downstream reporting tools.
When to use this
Extract UTM parameters on every landing page to capture traffic attribution at the session level, then forward them with conversion events so your analytics data ties revenue back to the correct campaign.
Examples
Collect and store UTM parameters on page load
// On page load — store UTM params if present const url = new URL(window.location.href); const UTM_KEYS = ["utm_source","utm_medium","utm_campaign","utm_term","utm_content"]; const stored = JSON.parse(sessionStorage.getItem("utm") || "{}"); const fresh = {}; for (const k of UTM_KEYS) { const v = url.searchParams.get(k); if (v) fresh[k] = v; } if (Object.keys(fresh).length) { sessionStorage.setItem("utm", JSON.stringify(fresh)); }
Read UTM parameters in a Python analytics handler
from urllib.parse import parse_qs, urlparse def extract_utm(raw_url: str) -> dict: query = urlparse(raw_url).query params = parse_qs(query) utm_keys = ["utm_source","utm_medium","utm_campaign","utm_term","utm_content"] return {k: params[k][0] for k in utm_keys if k in params}
- 1.
"UTM parameters," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/UTM_parameters
- 2.
Mozilla Developer Network, "URLSearchParams," developer.mozilla.org, accessed June 2026. https://developer.mozilla.org/en-US/docs/Web/API/URLSearchParams
- 3.
"General Data Protection Regulation," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/General_Data_Protection_Regulation
- 4.
Google, "Custom campaigns – Analytics helps," support.google.com, accessed June 2026. https://support.google.com/analytics/answer/10917952
- 5.
WebKit, "Tracking Prevention in WebKit," webkit.org, accessed June 2026. https://webkit.org/tracking-prevention/
utm_source identifies the traffic origin (e.g., 'google', 'newsletter'). utm_medium is the marketing channel ('cpc', 'email', 'organic'). utm_campaign is the specific campaign name. utm_term captures paid search keywords. utm_content differentiates ads or links in split tests. The first three are required for meaningful attribution; utm_term and utm_content are optional.
On the first page load, save the UTM parameters from the URL to sessionStorage. On subsequent pages, read from sessionStorage rather than the URL, because the URL may not carry UTM parameters for navigations that originate from your own site. When a conversion event fires, attach the stored UTM values to the event payload.
UTM parameters in the URL are visible to search engines. Google ignores utm_ prefixed parameters for canonicalization purposes, but other search engines may not. For SEO-sensitive pages, use rel='canonical' to point to the clean URL without UTM parameters, and configure your server or CDN to strip UTM parameters before caching responses.
utm_term is used for paid search to capture the keyword that triggered the ad. utm_content differentiates between multiple links or ads in the same campaign, for example, two different banner sizes in the same email. Both are optional and mainly used for granular attribution in A/B tests and paid campaigns.
In JavaScript: const url = new URL('https://example.com/landing'); url.searchParams.set('utm_source', 'newsletter'); url.searchParams.set('utm_medium', 'email'); url.searchParams.set('utm_campaign', 'spring-2026'); console.log(url.href). URLSearchParams encodes the values automatically, so no manual percent-encoding is needed. CapyToolkit offers a URL builder tool that constructs UTM-tagged URLs with proper encoding.