Scrub PII From CSV, JSON, Logs, Database Dumps and Git History

Remove emails, IPs, card numbers and credentials from CSV exports, JSON payloads, application logs, SQL dumps and git log output in your browser, while the file keeps its structure.

Scrub PII From CSV, JSON, Logs, Database Dumps and Git History

Exports are where personal data piles up. One CRM export, one day of access logs or one pg_dump can hold thousands of email addresses, phone numbers and IP addresses, and the moment you attach that file to a support ticket or paste it into an AI tool, every one of those values goes with it. Scrubbing the text first replaces each value with a numbered token while the commas, braces, log prefixes and SQL statements stay exactly where they were.

The scrubber reads text, so every format on this page follows the same three moves: get the text out of the file, scrub it in one paste, and check that the result still parses. What changes is where the personal data hides in each format. CSV comes first because it is the format most exports default to, then JSON, application logs, database dumps and git history.

Before you share an exported file

  • One paste per related set values that repeat across rows or tables get the same token only within one scrub
  • Structure check open the scrubbed CSV or JSON in its usual reader before you send it
  • Free-text columns names and street addresses in notes or comment fields need a manual pass
  • Variables file keep it apart from the scrubbed file, since together they rebuild the original

Opens the PII Scrubber with this page's checklist shown at the top of the tool.

Open in the tool →

Getting a file into the scrubber and back out intact

The scrubber works on pasted text, so binary formats need one extra step and text formats need none. That difference, plus a few habits about how much to paste at once and how to check the result, decides whether the scrubbed file is still useful to the person who receives it.

Copy the text, not the file

CSV, JSON, NDJSON, log files, SQL dumps and git log output are already plain text: open them in an editor, select all and paste. Spreadsheets and database clients need a copy from the grid or an export to CSV first, and pasted spreadsheet cells arrive tab-separated, which keeps the column boundaries intact in the scrubbed output.

For a file too large to paste comfortably, split it at a record boundary, such as a line in a log or a complete INSERT statement in a dump, never in the middle of a record. A record cut in half can hide a value from the patterns, because an email address or a card number split across two pastes no longer matches as a whole.

Keep related records in one paste

Within one scrub, a repeated value always gets the same token, so a customer email that appears in a users table and an orders table becomes the same [EMAIL_1] in both. That consistency is what lets a developer or an analyst still join, group and count the scrubbed data without ever seeing the real values behind the tokens.

Splitting related records across two scrubs breaks the link, because each run numbers its tokens from 1 again. When you must split, keep each table, or each log window, together in one run, and keep the variables file from each run with the part it belongs to so every reply can still be restored.

Check that the result still parses

Tokens contain no commas, quotes or braces, so a scrubbed file keeps its syntax. In CSV, a field that contains a comma, a double quote or a line break has to be enclosed in double quotes, and the scrubber leaves those quotes untouched.1 In JSON, a string value sits between double quotes, so a token that replaces an email inside a string is still a valid string.2 Open the result in the program that will read it before you send it, which catches the rare paste that was cut short.

Sources
  1. 1.

    Y. Shafranovich, "Common Format and MIME Type for Comma-Separated Values (CSV) Files," RFC 4180, IETF, October 2005. https://www.rfc-editor.org/rfc/rfc4180.html

  2. 2.

    T. Bray, "The JavaScript Object Notation (JSON) Data Interchange Format," RFC 8259, IETF, December 2017. https://www.rfc-editor.org/rfc/rfc8259.html

Scrub PII from CSV Files Before Sharing or Uploading

CSV files carry more PII than any other export format. CRM exports, HR spreadsheets, billing records, and support ticket dumps all default to CSV, and each row can contain a dozen identifiable fields such as names, emails, phone numbers, addresses, and account numbers.1 Sharing or uploading a CSV without scrubbing means every consumer of that file inherits the full data liability.

The risk multiplies when the destination is an AI tool. CSV content uploaded to analytics platforms, pasted into LLMs for processing, or shared with data vendors exposes every row to systems outside your control.2 Scrubbing PII from the CSV text before it leaves your machine removes the identifiable values while preserving the structure and non-sensitive content that downstream workflows need.

Opens the PII Scrubber with the value from this section already filled in.

Open in the tool →

What PII appears in CSV files

Customer CSVs typically contain email addresses, phone numbers, and physical addresses. HR CSVs add Social Security Numbers, salary figures, bank account details, and IBAN payment references. Billing CSVs include credit card numbers, card types, and associated email addresses. Building on this, many CRM exports include internal IDs, account numbers, and customer-facing domain names that may qualify as indirect identifiers under GDPR or CCPA. The scrubber detects all 22 standard PII types and replaces each with a numbered token so structure is maintained while values are hidden.

You should assume that any CSV exported from a business system contains PII until you verify otherwise. Even exports that appear structural, such as a product catalog or inventory list, may contain supplier contact emails or internal cost codes that qualify as business-sensitive data. The scrubber processes every line of the CSV in a single pass, so there is no performance penalty for scrubbing a file that turns out to contain no PII. The cost of one unnecessary scrub is a few seconds. The cost of one missed credential in a shared CSV can be a compliance incident.3

Scrubbing CSV text in practice

Open the CSV in a text editor or copy its content directly from a spreadsheet application (Excel, Google Sheets, LibreOffice). Paste the raw CSV text into the scrubber. The tool processes every line simultaneously, detecting emails, phone numbers, IBANs, and credit card patterns across all rows. Consequently, a 500-row export is scrubbed as fast as a 5-row sample. Download the variables file to maintain the mapping, then paste or save the scrubbed CSV for sharing with your AI tool, vendor, or internal data team.

When you copy from a spreadsheet application, the clipboard typically contains tab-separated values rather than comma-separated values. The scrubber handles both formats equally well because it processes the text as a flat string, not as a parsed table.4 After scrubbing, the output uses the same delimiter as the input. If you pasted tab-separated data, the scrubbed output is tab-separated and can be pasted directly back into your spreadsheet without reformatting. This makes the scrub-and-restore workflow seamless for spreadsheet-heavy teams.

Maintaining CSV structure after scrubbing

The scrubber replaces only the sensitive values. Commas, column headers, non-PII fields, and numeric values remain intact. A CSV with headers like email,phone,account_id,plan becomes the same header row followed by rows where email and phone are tokenized but account_id and plan remain readable. Yet this structure preservation depends on the scrubber detecting values correctly, so review the variables file after scrubbing to confirm every detected item was expected and check that non-PII numbers like account IDs were not flagged as SSNs or credit card numbers.

After scrubbing, validate the output by spot-checking a few rows against the original. Confirm that the column count per row has not changed, that quoted fields containing commas are still properly quoted, and that the header row is intact. For CSV files with complex quoting (fields that contain commas or newlines within quotes), the scrubber preserves the quoting structure because it operates on the raw text without parsing the CSV grammar.4 A quick visual check of the first and last few rows is usually sufficient to confirm structural integrity.

Common CSV export sources and the PII they contain

CRM platforms generate the most PII-dense CSV exports in most organizations. A Salesforce CSV report that includes Account and Contact objects can contain email addresses, phone numbers, mailing addresses, and company-associated names in a single export file. HubSpot contact exports add lifecycle stage, original traffic source, and associated deal values alongside contact identifiers. Marketo form submission exports include user-submitted form fields that may contain free-text entries with personal information beyond what the form was originally designed to collect.

HRIS platforms produce CSV exports with the highest regulatory sensitivity. Workday and BambooHR payroll exports include employee SSNs or national tax IDs for payroll processing, IBAN numbers for direct deposit, personal email addresses collected during onboarding, and salary data in numeric form. These files require scrubbing before any use with AI tools, even for structural tasks like reformatting column headers or generating payroll summaries.

Filtering which columns to scrub in large CSV exports

For large CSV exports where only specific columns contain PII, copy the PII-containing columns rather than the full file. In Excel or Google Sheets, select the email, phone, and SSN columns using Ctrl+Click on column headers, copy, and paste into the scrubber. Download the variables file and the scrubbed column values, then replace the original column values in your spreadsheet with the scrubbed version before any sharing or analysis step. This targeted approach reduces variables file size and makes the token-to-value mapping easier to review than scrubbing a full wide-table export.

Using the variables file to reconstruct and validate scrubbed CSV data

The variables file the scrubber generates maps each token to its original value in a downloadable JSON or CSV format. For CSV workflows where you later need to restore AI-generated analysis results back to real values, the Restore tab accepts both the variables file and the text containing tokens. Pasting an AI-generated CSV analysis response into the Restore tab and uploading the variables file replaces every token with the original value in one pass, producing a final output that reads with real names and contact details while never having transmitted them to the AI provider.

Validating the variables file after a large scrub is a critical quality step. Open the downloaded JSON and scan the detected entries to confirm that every token maps to an expected value and that no legitimate data fields (such as numeric account IDs or order numbers) were incorrectly tokenized. Credit card detection uses Luhn algorithm validation to reduce false positives, but 16-digit account numbers that happen to pass the Luhn check may still be captured.5 Correcting these in the variables file before restoration prevents incorrect substitutions in the final output.

Maintaining referential integrity across multiple scrubbed CSV files

When sharing multiple related CSV files with an AI tool (for example, a customer file and an orders file that share a customer email as the join key), scrub both files in the same browser session without closing the tab between pastes.6 Within a single browser session, the scrubber assigns the same token to the same value across all pastes. An email that appears in both files receives [EMAIL_1] in both, preserving the join key relationship in the tokenized output so the AI can correctly associate orders with customers across tables.

Sharing scrubbed CSV data with AI analytics tools

ChatGPT's Advanced Data Analysis feature accepts CSV files as uploads and runs Python analysis on them.7 Uploading a scrubbed CSV provides the full structural data that analysis requires: column names, row counts, value distributions, and aggregations all work correctly on tokenized data. The AI can generate pivot tables, calculate summary statistics, and produce charts from the scrubbed file. Only value-based lookups that require matching real email addresses or phone numbers against an external reference fail with tokens, and these can be handled post-analysis using the variables file to re-map results.

Pandas-based AI prompts that ask the AI to write Python code for CSV manipulation work fully with tokenized CSV. Generated Pandas code operates on the token strings as if they were real values; column filtering, merging, groupby operations, and string pattern matching on non-PII fields all function correctly. After the AI generates the analysis code, run it on the original CSV locally (not via the AI) when real value output is needed, since running the same code on real data locally avoids the tokenization round-trip for workflows where the final output requires original values.

Browser performance considerations for very large CSV pastes

The scrubber processes text in-browser with no character limit imposed by the tool. For CSV files with hundreds of thousands of rows, pasting the full content may cause brief processing delays in browsers with limited JavaScript heap. In Chrome and Edge, increasing the heap size using the js-flags launch flag set to max-old-space-size=4096 provides more headroom for very large pastes. For CSV files above 50,000 rows, processing in chunks of 10,000 rows per paste keeps each processing step fast and keeps the variables file per chunk small enough to review manually.

The chunking approach keeps the variables file reviewable even for very large exports, because each paste produces its own mapping. Since CapyToolkit processes the CSV entirely in your browser with no upload, you can split a 100,000-row file into manageable pieces without ever sending a row of customer data to any server.

When to use this

Use this before sharing any CSV export with an AI tool, external vendor, or analytics platform. Whenever the file contains customer emails, phone numbers, payment data, or employee records, flag payment fields before an external CSV share so downstream recipients never see the raw values.

Examples

CRM customer export

Before
name,email,phone
John Smith,[email protected],+1-415-555-0100
Sarah Lee,[email protected],(312) 555-0182
After
name,email,phone
John Smith,[EMAIL_1],[PHONE_1]
Sarah Lee,[EMAIL_2],[PHONE_2]

Each unique value gets its own numbered token, so [EMAIL_1] and [EMAIL_2] restore independently.

HR payroll CSV with employee SSNs

Before
name,ssn,iban
Alice Brown,123-45-6789,GB29NWBK60161331926819
Bob Davis,987-65-4321,DE89370400440532013000
After
name,ssn,iban
Alice Brown,[SSN_1],[IBAN_1]
Bob Davis,[SSN_2],[IBAN_2]
Sources
  1. 1.

    Mark Tressler, "Export to CSV – The Widespread Security Risk in Plain Sight," rowzero.com, September 2024. https://rowzero.com/blog/export-to-csv-security-risk

  2. 2.

    Lokke Moerel and Marijn Storm, "Data subjects as controllers: Violation of GDPR's fairness principle," iapp.org, September 2024. https://iapp.org/news/a/data-subjects-as-controllers-violation-of-gdpr-s-fairness-principle

  3. 3.

    Microsoft, "Supported entities," microsoft.github.io, accessed June 2026. https://microsoft.github.io/presidio/supported_entities/

  4. 4.

    Y. Shafranovich, "Common Format and MIME Type for Comma-Separated Values (CSV) Files," RFC 4180, IETF, October 2005. https://www.rfc-editor.org/rfc/rfc4180.html

  5. 5.

    Stripe, "What is the Luhn algorithm and how does it work?," stripe.com, accessed June 2026. https://stripe.com/resources/more/how-to-use-the-luhn-algorithm-a-guide-in-applications-for-businesses

  6. 6.

    Skyflow, "Tokenization," skyflow.com, accessed June 2026. https://docs.skyflow.com/docs/tokenization/overview

  7. 7.

    OpenAI, "Data analysis with ChatGPT," help.openai.com, June 2026. https://help.openai.com/en/articles/8437071-data-analysis-with-chatgpt

FAQ

Yes. There is no row limit. Paste the full CSV text and the scrubber processes every line in a single pass. For very large files (100,000+ rows), copy and paste in chunks if your browser slows down; each chunk produces an independent variables file.

Yes. Only the detected PII values are replaced. Column names, non-PII fields, numeric account IDs, and structural characters (commas, quotes, line breaks) are preserved exactly as they were.

Review the variables file after scrubbing. Credit card detection requires a Luhn-valid 13–16 digit number, so random numeric IDs will generally not match. If one does, note the mapping and avoid restoring that token in the final output.

The scrubber processes text, not binary Excel files. Copy the cell content from Excel (Ctrl+A, Ctrl+C while in a sheet) and paste it into the scrubber. The tab-separated output preserves cell boundaries.

Yes. The scrubbed CSV is valid CSV with tokens replacing PII values. CapyToolkit places no restrictions on how you use the scrubbed output: the tokens serve as stable anonymized identifiers that can be used for analysis without exposing real values.

Scrub PII from JSON Logs and API Payloads

JSON is the default format for API logs, webhook payloads, and database document exports.1 A single API response can contain dozens of PII fields such as user IDs linked to emails, nested address objects, credit card metadata, and session tokens, often deeply nested in ways that make manual review impractical.

When you paste JSON payloads into AI tools for debugging, analysis, or transformation assistance, every nested field travels with the paste. The scrubber detects sensitive values regardless of nesting depth, replacing them with tokens while preserving the full JSON structure that makes the payload meaningful for the downstream task.

Opens the PII Scrubber with the value from this section already filled in.

Open in the tool →

PII in JSON structures

API responses nest PII at multiple levels. A user object might contain an email at the top level and a phone number nested under the address object, with a session JWT embedded in a separate metadata field. Consequently, a manual inspection of a 50-field JSON document is both time-consuming and error-prone. The scrubber processes the full text, detecting email patterns, JWT headers (eyJ prefix), IPv4 addresses, and credential formats across every field regardless of nesting.2 JSON structure, including keys, brackets, and non-PII values, is preserved after tokenization.

You should pay special attention to nested objects that contain both identifiers and authentication data. A single API response might include a user profile with an email address, a payment method object with a credit card number, and a session object with a JWT token. All three are PII, and all three are detected by the scrubber regardless of how deeply they are nested. The scrubber does not need to understand the JSON schema; it scans the raw text for pattern matches, so even unconventional nesting structures or dynamically generated field names do not prevent detection.

Common JSON sources that contain sensitive data

API logs are the most frequent source: request logs include caller IPs and authorization headers; response logs include user data and session tokens. Webhook payloads from payment providers include card metadata. Database document exports from MongoDB or Firestore include user profile documents with email addresses, phone numbers, and address fields. Furthermore, CI/CD pipeline outputs, error tracking payloads, and observability traces (OpenTelemetry, Datadog) frequently include request context with internal IPs and auth tokens.

You should treat any JSON that has passed through a production system as potentially containing PII.3 Error tracking payloads from Sentry, Datadog, or Rollbar frequently include request context with user IPs, session tokens, and authenticated user email addresses. Observability traces from OpenTelemetry or AWS X-Ray include service-to-service call metadata that may contain internal hostnames and API keys. Before pasting any of these JSON sources into an AI tool for analysis, run them through the scrubber to catch the embedded identifiers that are easy to overlook in deeply nested trace data.

Preserving JSON validity after scrubbing

The scrubber performs string replacement, not JSON parsing. It replaces raw string values that match PII patterns. Because it targets values rather than structure, keys, brackets, commas, and colons are untouched. Yet some edge cases, such as a URL string that contains an email address or a description field that includes a phone number in prose, will be tokenized in place. The resulting JSON remains syntactically valid and can be parsed, diffed, or imported into another tool after scrubbing.

After scrubbing, you can validate the JSON by pasting it into any JSON parser or validator. The token format [EMAIL_1] is a valid JSON string value, so the scrubbed output parses identically to the original from a structural perspective.4 For workflows where you need to diff the scrubbed JSON against the original (for example, to verify that only PII fields were changed), a text diff will show exactly which values were replaced. Non-PII fields, numeric values, boolean values, and null values are all preserved exactly as they were in the original.

Handling arrays and high-volume JSON payloads

JSON arrays of objects are processed by the scrubber in a single pass across the full text. A 500-element user array containing an email and phone number in each object is scrubbed as fast as a single-object response, because the scrubber performs text pattern matching rather than JSON tree traversal. Nesting depth has no effect on detection accuracy: an email address at five levels of nesting is detected the same way as one at the top level.

For very large JSON payloads (API log exports with thousands of entries), paste the content in segments if your browser slows down during processing. Each segment produces an independent variables file; download and store all segment variables files together so the complete token-to-value mapping is available when you need to restore. Within each segment, the same value always maps to the same token, so [EMAIL_1] in segment one and [EMAIL_1] in segment two represent the same original email address only if the segments were scrubbed in separate passes from separate source texts.

Minified JSON and formatted JSON are both supported

The scrubber processes raw text regardless of formatting, which means you do not need to reformat your JSON before pasting it in. Minified JSON (all on one line with no whitespace) and pretty-printed JSON (indented with newlines) both work correctly because the pattern matching targets values within the string, not structural characters. Neither format produces different detection results, so you can paste API responses directly from your browser's network inspector or from a log file without any preprocessing step.

JWT tokens embedded in JSON API responses

JWT tokens appear frequently in JSON as the value of fields named token, access_token, id_token, authorization, or session. The three-segment eyJ prefix detection the scrubber uses catches all standard JWTs regardless of which JSON field name contains them. Each JWT is replaced with a distinct [JWT_N] token, so multiple JWTs in the same payload are tracked separately in the variables file.

JWTs embed claims in their payload segment. When a JWT contains a user's email address, user ID, or organizational claims, those values are encoded inside the token itself. The scrubber tokenizes the full JWT string, which effectively removes all the claims encoded in it in a single step.2 If you need to share the claims structure (to show a JWT's payload to an AI tool for schema review), decode locally, then scrub the JWT claims using a library or jwt.io before pasting, and share only the scrubbed claim structure.

Authorization headers in JSON-formatted HTTP request logs

HTTP request logs formatted as JSON often include an Authorization header field containing a Bearer prefix followed by a JWT or API key. The scrubber detects both the JWT format and the API key patterns in this field regardless of the Bearer prefix. After scrubbing, the log retains the Authorization key and the Bearer prefix with the token replaced by a placeholder, preserving the structural context that shows an authenticated request was made.

NDJSON log output and document database exports

NDJSON (Newline Delimited JSON) is the standard output format for log aggregation tools, MongoDB mongoexport, Elasticsearch exports, and Datadog log archives. Each line in an NDJSON file is a complete JSON object. The scrubber processes NDJSON as plain text: paste the NDJSON content directly and each line's values are detected and tokenized independently. The output is valid NDJSON with tokens replacing sensitive values, preserving the one-document-per-line format that downstream tools expect.

MongoDB mongoexport produces NDJSON by default when exporting a collection.5 A user collection export contains one document per line with email, phone, address, and account fields as JSON object values. Pasting a mongoexport output into the scrubber tokenizes all PII across every document in a single pass, giving you a scrubbed export suitable for sharing with a developer or AI analysis tool without transmitting any production user records.

Differencing scrubbed JSON for AI-assisted schema review

Pasting two versions of the same JSON schema (before and after a field addition, for example) into an AI tool for comparison is a common developer workflow. Scrubbing both versions before pasting ensures that example values in the schema do not expose real data. JSON schema comparison requires structural information (key names, value types, array shapes) but never requires the actual example values.6 The scrubber removes the values and preserves the structure that makes schema comparison useful.

Because the scrubber preserves keys and structure while replacing only values, the diff shows exactly which fields changed without revealing the sample data behind them. CapyToolkit runs the comparison preparation locally in your browser, so the JSON never leaves your machine during scrubbing and the variables file is the only artifact that maps the tokens back to the real schema examples.

When to use this

Use this before pasting any API payload, webhook body, or JSON document export into an AI tool when the payload contains user data, authentication tokens, or internal service addresses.

Examples

API user response with nested PII

Before
{"id": 42, "email": "[email protected]", "phone": "+14155550100", "token": "eyJhbGciOiJSUzI1NiJ9.eyJ1c2VyX2lkIjoiMTIzIn0.sig"}
After
{"id": 42, "email": "[EMAIL_1]", "phone": "[PHONE_1]", "token": "[JWT_1]"}

JSON structure and non-PII values (id: 42) are preserved. Paste the scrubbed JSON into your AI tool for transformation help.

Webhook payload with payment metadata

Before
{"event": "charge.succeeded", "customer_email": "[email protected]", "card_last4": "4242", "ip": "203.0.113.5"}
After
{"event": "charge.succeeded", "customer_email": "[EMAIL_1]", "card_last4": "4242", "ip": "[IP_1]"}
Sources
  1. 1.

    Daniele Molteni, "Landscape of API Traffic," blog.cloudflare.com, January 2022. https://blog.cloudflare.com/landscape-of-api-traffic/

  2. 2.

    M. Jones, J. Bradley, and N. Sakimura, "JSON Web Token (JWT)," RFC 7519, IETF, May 2015. https://www.rfc-editor.org/rfc/rfc7519.html

  3. 3.

    Sentry, "User Interface," develop.sentry.dev, accessed June 2026. https://develop.sentry.dev/sdk/foundations/envelopes/event-payloads/user/

  4. 4.

    IETF, "The JavaScript Object Notation (JSON) Data Interchange Format," RFC 8259, IETF, December 2017. https://www.rfc-editor.org/rfc/rfc8259.html

  5. 5.

    MongoDB, "mongoexport," mongodb.com, accessed June 2026. https://www.mongodb.com/docs/database-tools/mongoexport/

  6. 6.

    Andrey Vit, "json-diff," npmjs.com, accessed June 2026. https://www.npmjs.com/package/json-diff

FAQ

Yes. The scrubber processes the full text of the JSON, so nested values are detected and tokenized regardless of depth. It works on the raw string representation, which is the same text you see when you copy a JSON response.

Yes. Because the scrubber replaces values within quoted strings, the resulting JSON is syntactically valid. Tokens like [EMAIL_1] are valid string values. The structure, keys, and non-PII fields are unchanged.

Yes. If your log file contains one JSON object per line (newline-delimited JSON / NDJSON), paste the full content and the scrubber processes every line. If it is a JSON array, paste the full array text.

JWT tokens (eyJ prefix, three base64-encoded segments), database connection string values, API keys in multiple formats, IP addresses, email addresses, credit card numbers, SSNs, phone numbers, and IBANs. The scrubber detects based on value patterns, not JSON field names.

Yes. After scrubbing, check the Variables section of the tool. Every detected value is listed with its token and original value. This doubles as a PII audit of your JSON payload.

Scrub PII from Application and Server Logs

Logs record everything your application does, including things it shouldn't. Request logs capture caller IPs and auth headers. Error logs include stack traces with internal hostnames. Access logs record user email addresses used as identifiers. Debug logs often include full request bodies with customer data embedded.

Sharing application logs for debugging, incident response, or vendor support tickets is one of the most common accidental PII disclosure paths. The scrubber removes the identifiable values from log text before it leaves your environment, letting you share the operational signal without the sensitive context.

Opens the PII Scrubber with the value from this section already filled in.

Open in the tool →

PII patterns in application logs

Web server access logs include client IP addresses on every line. Application logs at INFO and DEBUG level frequently include user identifiers, often email addresses used as usernames. Error logs catch full exceptions that include database connection strings in the message, internal hostnames in the stack trace, and auth tokens in HTTP header dumps.1 Consequently, a simple log grep for an error message can return a file that, pasted verbatim into a support ticket or AI debugging session, exposes dozens of sensitive values. The scrubber reduces a raw log to its operational signal.

You should assume that any production log file contains PII until you verify otherwise. Even logs from internal-only services can contain developer email addresses in authentication contexts, internal IP addresses in connection metadata, and API keys in configuration dumps. The scrubber processes log text line by line, detecting all 22 PII patterns regardless of which log format you use.2 A 10,000-line log file is scrubbed in the same time it takes to paste it, so there is no practical reason to skip the scrub step even for large log exports.

Log formats the scrubber handles

The scrubber works on plain text, so it handles any log format: Apache/Nginx combined log format (IP at the start of each line), JSON log output (structured fields), Python exception tracebacks (connection strings in the message), Node.js application logs (email identifiers in user objects), and CI/CD pipeline output (auth tokens in environment variable dumps). Furthermore, OpenTelemetry traces exported as text or NDJSON are processed the same way as any other structured text, with field-level tokenization and no structure loss.1

You do not need to reformat your logs before scrubbing. Paste them exactly as they appear in your log viewer, terminal, or log aggregation tool. The scrubber detects PII patterns in raw text, so formatting differences such as colored output, timestamp prefixes, or structured field labels do not affect detection accuracy. If your logs include ANSI color codes from terminal output, the scrubber processes them as part of the text and detects PII patterns around them without any preprocessing needed.

Safe log sharing practices

After scrubbing, the log file is safe to share with vendors, AI debugging tools, or external support channels. The variables file maps each token back to its original value, so your internal team can cross-reference the shared log against the real values without reconstructing them manually. Building on this, scrubbed logs can safely be stored in shared incident management systems like Jira or GitHub Issues where broader team access is expected. Yet the variables file itself should stay restricted, and you should treat it with the same access control as the original log.

You should establish a clear retention policy for variables files generated from log scrubbing. The variables file is the key that reconstructs the original log content, so it carries the same sensitivity classification as the original log. For incident response workflows, store the variables file in the same secure location where you would store the original log export. When the incident is resolved and the log is no longer needed, delete the variables file alongside the log export. Treating the variables file as a transient artifact rather than a permanent record keeps your data lifecycle consistent.

Structured logging frameworks and PII propagation

Structured logging frameworks (Winston, Bunyan, Logstash, Datadog Agent) emit log lines as JSON objects with consistent field schemas.3 These frameworks add session context to every log line for a given request: a middleware that logs the authenticated user's email with the first event of a session propagates that email to every subsequent log line in the same session. A single user action can generate dozens of log lines, each inheriting the email address as a field value.

The scrubber processes structured JSON log output (NDJSON) the same way it processes any text: paste the log export and all field values matching PII patterns are tokenized across every line. A session with 200 log lines sharing the same user email produces a single [EMAIL_1] token that represents the email across all 200 lines in the output. The variables file stores the mapping once, keeping the token-to-value reference compact regardless of how many log lines contained the original value.

Log enrichment pipelines and PII amplification

Log enrichment pipelines (used in ELK Stack, Splunk, and Datadog) add geolocation data, user profile fields, and organization metadata to raw log lines as they pass through the pipeline.4 An access log line containing only a client IP address may exit the enrichment pipeline with that IP resolved to a user email, name, and account ID. Logs exported after enrichment carry significantly more PII per line than raw logs. Identify whether your logs have been enriched before scrubbing, and if so, expect more PII fields per line and a larger variables file.

Shipping log excerpts to vendor support and AI tools

Vendor support tickets for infrastructure products (Datadog, AWS, Cloudflare, PagerDuty) routinely require log excerpts to diagnose issues. These platforms' support teams are third parties under most data governance frameworks, which means sharing logs containing customer IP addresses or user emails with vendor support requires either a data processing agreement or removal of the PII from the excerpt.

Scrubbing the relevant log lines before attaching them to a support ticket satisfies both requirements simultaneously: the vendor receives the operational signal they need to diagnose the issue (error codes, timestamps, response codes, stack trace structure) without receiving any customer identifiers. After the ticket is resolved, you can restore the real values from the variables file if you need to reconstruct the full incident timeline internally.

Incident response workflows and scrubbed log handoffs

During an active incident, speed is critical and scrubbing can feel like a delay. Establish a pre-scrub step as part of your standard handoff procedure rather than an ad-hoc decision under pressure. A runbook section titled "Before sharing logs externally" that references the scrubber URL makes the step routine rather than optional. Scrub a 50-line excerpt under time pressure in under 30 seconds; building the habit eliminates a category of PII exposure that otherwise occurs most frequently during the high-pressure moments when careful review is hardest.

Less-obvious PII patterns in application logs

Beyond IP addresses and email addresses, application logs contain PII in less-obvious forms. HTTP Authorization headers logged in request middleware include Bearer tokens and Basic auth credentials encoded in base64. User-agent strings from misconfigured client libraries occasionally include email addresses or user IDs. Full request body logging at DEBUG level records whatever the user submitted: a search query, a form field, or a free-text comment that may contain names, contact information, or health details.

Review your logging configuration for any middleware or interceptor that logs request headers, full request bodies, or raw query strings. These are the sources of the least-expected PII in your logs. For logs from these sources, the scrubber catches the structured patterns (IPs, emails, auth tokens) but cannot catch free-form user-submitted text that does not match a known pattern. For high-sensitivity applications, review the scrubbed output manually for any user-submitted prose fields that may contain personal information beyond what the pattern matching caught.

Authorization token logging and the JWT detection scope

Bearer tokens in Authorization header logs often contain JWTs with user identity claims embedded in the payload. The scrubber's JWT detection covers the eyJ prefix plus the complete three-segment token structure, so JWTs logged in any field are tokenized regardless of how they appear in the log line.5 Basic auth credentials logged as Authorization: Basic <base64> are not in the default detection set because base64 strings without the JWT prefix do not match the JWT pattern; remove or redact Basic auth header values manually if they appear in your logs.

Because the scrubber runs in your browser with no server upload, the log text you scrub never leaves your machine during processing. For the Basic auth values the JWT pattern misses, a quick manual redaction before pasting closes the remaining gap, and the variables file then holds the only mapping back to the real credentials for safe restoration.

When to use this

Use this before sharing application logs, access logs, or error traces with vendors, support teams, AI debugging tools, or any system outside your internal network.

Examples

Nginx access log with client IPs

Before
192.168.1.10 - [email protected] [01/Jun/2026] "GET /api/data HTTP/1.1" 200
10.0.0.5 - [email protected] [01/Jun/2026] "POST /api/auth HTTP/1.1" 200
After
[IP_1] - [EMAIL_1] [01/Jun/2026] "GET /api/data HTTP/1.1" 200
[IP_2] - [EMAIL_2] [01/Jun/2026] "POST /api/auth HTTP/1.1" 200

Python traceback with connection string

Before
psycopg2.OperationalError: could not connect to server: postgres://admin:[email protected]/main
  File "db.py", line 42, in connect
After
psycopg2.OperationalError: could not connect to server: [DBURL_1]
  File "db.py", line 42, in connect

The stack trace context (file name and line number) is preserved while the connection string is tokenized.

Sources
  1. 1.

    Apache HTTP Server, "Log Files," httpd.apache.org, accessed June 2026. https://httpd.apache.org/docs/2.4/logs.html

  2. 2.

    Microsoft, "Supported entities," microsoft.github.io, accessed June 2026. https://microsoft.github.io/presidio/supported_entities/

  3. 3.

    Trent M. node-bunyan, "log.child," github.com, accessed June 2026. https://github.com/trentm/node-bunyan/blob/master/README.md

  4. 4.

    Datadog, "Lookup Processor," docs.datadoghq.com, accessed June 2026. https://docs.datadoghq.com/logs/log_configuration/processors/lookup_processor/

  5. 5.

    M. Jones, J. Bradley, and N. Sakimura, "JSON Web Token (JWT)," RFC 7519, IETF, May 2015. https://www.rfc-editor.org/rfc/rfc7519.html

FAQ

Yes. The scrubber processes the full pasted text as a block. Multi-line entries such as Java exception stacks or Python tracebacks are processed line by line. Each unique sensitive value gets one token regardless of how many times it appears.

Each unique value gets its own token. If 192.168.1.10 appears 500 times, every occurrence is replaced with [IP_1]. The variables file stores the mapping once, not 500 entries.

Yes. That is the primary use case for this page. Paste the relevant log section into the scrubber, copy the cleaned output, and paste it into your AI tool. The AI gets the operational signal without any identifying values.

Yes. JSON log output is text. Paste it into the scrubber and the pattern matching runs over the full string, detecting values inside JSON fields without affecting the surrounding structure.

The scrubber detects PII based on patterns, not field names. If a user submitted their email as free text and it was logged, the scrubber will detect and tokenize it regardless of how it ended up in the log. CapyToolkit processes everything locally in your browser, so your log data never leaves your machine during scrubbing.

Scrub PII from Database Dumps Before Sharing

A database dump is a full copy of your production data. pg_dump, mysqldump, and mongodump output files contain every row of every table, including all the PII your application has ever stored. Sharing a database dump for debugging, developer onboarding, or vendor analysis is one of the highest-risk data operations a team can perform.1

Scrubbing a database dump before sharing is a practical way to reduce that risk without rebuilding your development database infrastructure. Paste the relevant tables or query results into the scrubber to replace PII with tokens, then share the anonymized version with the recipient who needs the data structure, not the real values.

Opens the PII Scrubber with the value from this section already filled in.

Open in the tool →

PII in database exports

User tables contain the densest concentration of PII: email addresses, phone numbers, hashed or plaintext passwords, addresses, and date-of-birth fields. Order tables add credit card references, shipping addresses, and transaction IDs. For SaaS applications, subscription tables contain billing email addresses and payment method references. Consequently, a partial dump of just the users and orders tables from a mid-sized application can contain millions of identifiable records. Scrubbing the key fields before sharing reduces this to structural data that developers need for schema analysis or query debugging.

You should treat every database dump as a high-risk data artifact, regardless of which tables it contains. Even a dump of what appears to be a configuration or lookup table may contain admin email addresses, API keys stored as configuration values, or internal hostnames in endpoint URLs. The scrubber detects all 22 PII patterns in SQL text, including values inside quoted strings, comments, and stored procedure definitions. Before sharing any dump, scrub the PII-heavy tables and spot-check the remaining tables for embedded identifiers.2

How to scrub a SQL dump effectively

For a pg_dump or mysqldump output, focus on the INSERT INTO statements for PII-heavy tables: typically users, customers, orders, payments, and session tables. Copy those blocks into the scrubber separately. Email addresses, phone numbers, credit card patterns, SSNs, and IBANs are detected in SQL value lists. Building on this, IPv4 addresses and JWT tokens in session or auth tables are also caught. The scrubber processes quoted SQL strings correctly, with tokens replacing only the values inside quotes and the SQL syntax left intact.

You can scrub multiple related tables in the same browser session to preserve referential integrity. If the same email address appears in both the users table and the orders table, scrubbing them together in one session ensures both occurrences receive the same token. This means foreign key relationships survive the scrubbing process: a JOIN between the scrubbed users table and the scrubbed orders table produces correct results because [EMAIL_1] in the users table matches [EMAIL_1] in the orders table. Scrub all related tables in a single session rather than one at a time.3

Using a scrubbed dump for development

A common pattern is to scrub the PII fields and then share the cleaned dump with developers as a development seed database. Developers get a realistic data structure with correct table shapes, real row counts, and real relationship patterns, all without accessing any production PII. Yet the tokens are stable identifiers: [EMAIL_1] is always the same customer across every table it appears in, because the scrubber assigns a token to the unique value, not the occurrence. Referential integrity is preserved across tables for the same input.

You should validate the scrubbed dump before importing it into a development database. Run the scrubbed SQL against a test instance and verify that the row counts match the original, that foreign key constraints are satisfied, and that the schema DDL was not accidentally modified. The scrubber only touches quoted string values, so CREATE TABLE statements, indexes, and constraints pass through unchanged. A quick smoke test of the most common queries your application runs against the development database confirms that the scrubbed data behaves like the original from a query perspective.4

Prioritizing tables for scrubbing by PII density

Not every table in a database dump carries the same PII risk. Users, customers, contacts, orders, payments, and session tables hold the highest density of identifiable data and should always be scrubbed before sharing. Configuration tables, product catalog tables, and schema migration history tables typically contain no PII and do not require scrubbing. Schema DDL (CREATE TABLE and ALTER TABLE statements) contains column names and types but no data values, and is safe to share without scrubbing.

For a targeted partial dump shared for debugging purposes, target five tables, not the whole dump, copying only the INSERT blocks from your highest-risk tables and scrubbing those blocks specifically. Scrubbing INSERT blocks from five tables (users, customers, orders, payments, sessions) typically removes 90% or more of the PII exposure in a partial application database dump. The schema DDL and configuration tables can be shared separately without scrubbing, keeping the total volume of text you run through the scrubber manageable.

Query results versus full table exports

Single-query result exports (rows returned from a SELECT statement) often contain a subset of columns rather than the full table schema. Before scrubbing, identify which columns in the query result contain PII: email, phone, ssn, iban, ip_address, and token columns are the primary targets. The scrubber detects by value pattern rather than column name, so any PII-formatted value in any column is caught regardless of the column name used in your schema.

MongoDB and NoSQL document exports

MongoDB mongoexport produces NDJSON output: one document per line, with the full document structure as a JSON object. A user collection exported with mongoexport using the collection=users and out=users.ndjson flags generates one line per user document, each containing all the fields in that document. Email addresses, phone numbers, addresses, and account identifiers appear as JSON field values across every line in the export.

Paste the full mongoexport output into the scrubber. The pattern matching runs across every line simultaneously, tokenizing all PII field values in all documents in a single pass. The output is valid NDJSON with tokens replacing sensitive values, preserving the document structure and non-PII fields that developers need for schema analysis and query debugging. Import the scrubbed NDJSON into a development MongoDB instance using mongoimport with the collection flag for a fully functional development dataset with no production PII.5

Firestore, DynamoDB, and other document database exports

Firestore exports via gcloud firestore export produce a proprietary binary format that must be converted to JSON before scrubbing. Use the Firebase Admin SDK or community export tools to extract documents as JSON before passing them through the scrubber. DynamoDB exports via aws dynamodb scan with output set to json produce a JSON array of item objects; paste the JSON directly into the scrubber. The scrubber handles any JSON document format regardless of the underlying database that produced it.6

Development seed databases and PII propagation risk

Development and staging environments seeded from production dumps create long-lived copies of production PII that persist beyond the original sharing event. Developer workstations, local Docker containers, and shared development databases often have weaker access controls than production, widening the PII surface to everyone with development environment access. A single unscrubbed production dump seeded across ten developer machines creates ten independent copies of production customer data under ten different access control configurations.

Scrubbing the dump before seeding the development database eliminates PII from the development environment at the source. Developers receive a realistic data structure: correct table shapes, real row counts, real relationship patterns, and consistent foreign key references. The tokens serve as stable identifiers across tables for the session that produced them, so JOIN operations between scrubbed users and scrubbed orders tables work correctly in the development database. This approach reduces the regulatory classification of the development database to non-regulated status, removing it from the scope of PII-related compliance audits.

Automating scrubbing in the production-to-dev data pipeline

For teams that refresh development environments regularly from production snapshots, integrate the scrubbing step into the data pipeline rather than performing it manually each time. A pipeline stage that exports the target tables, scrubs the exported text, and imports the scrubbed data into the development database makes PII-free seeding automatic and consistent. The variables file produced by each pipeline run provides an audit record of what was replaced in each refresh cycle.

Wiring the scrub step into the pipeline means every refresh produces a development database with no production PII, by construction. CapyToolkit runs the scrubber locally in your browser, so the pipeline stage can call it without uploading dumps to a server, and the per-run variables file gives auditors a clear record of what was replaced in each cycle.

When to use this

Use this before sharing any table export, query result, or database dump with a developer, vendor, analytics team, or AI tool when the output contains customer or employee records.

Examples

SQL INSERT with user records

Before
INSERT INTO users (id, email, phone) VALUES (1, '[email protected]', '+14155550101'), (2, '[email protected]', '(312) 555-0188');
After
INSERT INTO users (id, email, phone) VALUES (1, '[EMAIL_1]', '[PHONE_1]'), (2, '[EMAIL_2]', '[PHONE_2]');

SQL syntax and non-PII values (id column) are preserved. The scrubbed INSERT is valid SQL that can be imported directly.

MongoDB document export with payment data

Before
{"_id": "u001", "email": "[email protected]", "card": "4111111111111111", "iban": "GB29NWBK60161331926819"}
After
{"_id": "u001", "email": "[EMAIL_1]", "card": "[CC_1]", "iban": "[IBAN_1]"}
Sources
  1. 1.

    PostgreSQL, "pg_dump," postgresql.org, accessed June 2026. https://www.postgresql.org/docs/current/app-pgdump.html

  2. 2.

    Microsoft, "Supported entities," microsoft.github.io, accessed June 2026. https://microsoft.github.io/presidio/supported_entities/

  3. 3.

    M. Jones, et al., "JSON Web Token (JWT)," RFC 7519, IETF, May 2015. https://www.rfc-editor.org/rfc/rfc7519.html

  4. 4.

    Morris Dworkin, "Recommendation for Block Cipher Modes of Operation: Methods for Format-Preserving Encryption," SP 800-38G, nist.gov, March 2016. https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-38G.pdf

  5. 5.

    MongoDB, "mongoexport," mongodb.com, accessed June 2026. https://www.mongodb.com/docs/database-tools/mongoexport/

  6. 6.

    Amazon Web Services, "scan," docs.aws.amazon.com, accessed June 2026. https://docs.aws.amazon.com/cli/latest/reference/dynamodb/scan.html

FAQ

The scrubber works on pasted text, not file uploads. For large dumps, copy the INSERT blocks for PII-heavy tables separately and scrub each section. Schema DDL (CREATE TABLE statements) typically contains no PII and does not need scrubbing.

Yes, for the same input. The scrubber assigns tokens based on unique values, so [EMAIL_1] always represents the same email address throughout the paste. If you scrub users and orders in the same paste, the email that appears in both tables gets the same token. CapyToolkit processes everything locally, so your production data stays in your browser.

Not intentionally. Bcrypt hashes (starting with $2b$) are not in the default detection patterns. They are treated as non-PII strings. If your dump includes plaintext passwords, they will be detected as generic API keys only if they match the sk- prefix pattern.

For one-off sharing and AI analysis, the scrubber is practical and fast. For systematic production-to-development database anonymization at scale, a dedicated masking tool (PostgreSQL anonymizer, Faker-based seeding) is more appropriate. The scrubber solves the spot-check and AI-sharing use case.

Yes. Download the variables file after scrubbing. Upload the file to the Restore tab along with the scrubbed dump text to reverse the tokenization. Keep the variables file as securely as the original dump.

Scrub Sensitive Data from Git History and Commit Messages

Git logs capture what your developers typed. Commit messages, git blame output, git log with body text, and git diff headers can all contain sensitive information such as email addresses used as commit authors, internal hostnames referenced in commit messages, and occasionally credentials accidentally included in diff context.

Sharing git history with vendors, AI tools, or external contributors exposes this accumulated context. The scrubber processes git log output as plain text, replacing email addresses, internal domain names, and credentials with tokens before the history leaves your repository environment.

Opens the PII Scrubber with the value from this section already filled in.

Open in the tool →

What sensitive data appears in git history

Git author fields contain developer email addresses in every commit. Commit message bodies frequently reference internal ticket URLs, internal server names, and coworker email addresses when describing changes. Git blame output adds these fields to every line of a file. Furthermore, accidental credential commits are common: developers occasionally commit .env files or config files with passwords, and even if later removed, these remain in git history until explicitly purged with git filter-repo or BFG Repo Cleaner1. The scrubber handles the sharing use case, not git history rewriting.

You should assume that any git log output you paste into an AI tool contains PII. Even a simple one-line-per-commit log includes author email addresses. A full-format log with commit bodies contains internal ticket references, coworker names, and server hostnames. Before pasting any git output into an AI tool for changelog generation, sprint analysis, or code archaeology, run it through the scrubber. The AI can analyze commit patterns, identify change categories, and generate release notes from tokenized output just as effectively as from the original.

Using git log output for AI analysis

Developers increasingly paste git log output into AI tools to analyze release changes, generate changelogs, or summarize what changed in a sprint. A git log one-line output contains commit hashes and messages but typically no PII. Yet git log with a format string that includes author email shows developer addresses in every entry. Consequently, before passing git log output to an AI tool, paste it through the scrubber to tokenize author emails and any internal references in commit message bodies.

You can scope your git log output before scrubbing to reduce the volume of text. Instead of dumping the entire repository history, use a date range or a branch range to limit the output to the commits relevant for your analysis. A focused log of 50 commits from the past sprint is faster to scrub and produces a more compact variables file than a full history dump. The AI receives the same structural information about what changed, and your developer identities stay private.

Vendor and AI tool sharing of git diffs

Code review and AI-assisted code analysis tools often accept git diff output. A git diff can include context lines (surrounding code), and those context lines may include email addresses in comments, internal hostnames in configuration values, and API key patterns in test fixtures. Building on this, git diff stat output shows only file names and change counts, which is safe to share2. Git diff unified output includes actual code lines and needs scrubbing if those lines touch config files or contain inline credentials. Run the diff text through the scrubber before sharing it externally.

You should scrub git diff output even when the changed code itself contains no credentials. The surrounding context lines in a diff often include import statements, configuration blocks, and comment headers that contain internal hostnames, developer emails, and API endpoint URLs. These context lines are included to help the AI understand the structural change, but they carry the PII that identifies your infrastructure. The scrubber tokenizes the identifiers in context lines while preserving the code structure that makes the diff meaningful for analysis.

Commit message bodies and internal references beyond author fields

Git commit message bodies contain more sensitive context than subject lines. Developers write detailed commit bodies that reference internal ticket URLs (Jira, Linear, GitHub Issues with internal project names), internal server names encountered during testing, and coworker email addresses when documenting co-authorship or review attribution. Merge commits generated by GitHub and GitLab include the full PR description3, which may reference internal infrastructure details and reviewer email addresses.

Running git log with the format string "%H%n%ae%n%s%n%b" and a commit range produces output that includes author email, subject, and full body text for each commit. This is the format that captures the most operational context for AI-assisted changelog generation, and it is also the format with the highest PII density. Scrub this output before passing it to any AI tool for changelog generation, sprint summary, or commit pattern analysis.

Co-author trailers and signed-off-by fields

Co-authored-by: and Signed-off-by: git trailers in commit message bodies contain email addresses4 that are part of the commit message text rather than the git author field. These trailers appear in every commit created via "Co-authored-by" workflows (common in pair programming and AI-assisted development), and they accumulate in git log output as developer emails from every contributor to a commit. The scrubber's email detection catches these trailer email addresses across all commit entries in the log output.

Git blame output and line-level developer attribution

Git blame annotates every line of a source file with the author's email address, the commit hash, the date, and the line content. Sharing git blame output reveals the complete attribution map of a file: which developer wrote each line, when, and in which commit. For code review tools, AI analysis tools, or vendor support conversations that require showing which part of the code introduced a bug, blame output combines sensitive developer attribution with potentially sensitive code content in a single export.

The scrubber tokenizes author email addresses in git blame output while preserving the commit hash, date, and line content. The recipient sees [EMAIL_1] for every line authored by the same developer, with distinct tokens for different authors. The structural attribution (this developer touched lines 40-60; that developer touched lines 61-80) remains intact while the developer identity is removed.

Exporting blame output for sharing

Generate blame output for sharing with git blame using the line-porcelain flag and a file path to get the machine-readable format with each field on its own line, making author emails straightforward to locate and verify after scrubbing. The porcelain format's author-mail field contains the email5; the scrubber catches this field value because it matches the email pattern regardless of the git blame output format. After scrubbing, the porcelain-format blame output remains structurally valid for downstream tools that parse it programmatically.

AI-assisted changelog generation from git history

AI tools are increasingly used to generate changelogs, release notes, and sprint summaries from git log output. The git log command with commit message bodies and author information provides the raw material for these summaries. Passing unscrubbed git log output to an AI tool exposes author email addresses and internal references from every commit in the specified range to the AI provider's infrastructure.

Scrub the git log output before passing it to any AI tool for changelog generation. The AI receives commit subjects, bodies, and structural metadata (dates, commit hashes) with author emails tokenized and internal references removed. The output it generates (a grouped changelog, a sprint summary, a release note) contains the same structural information derived from the commit history without including the developer identifiers. For commit bodies that reference internal ticket identifiers, the AI can still identify the nature of the change from the commit subject and code diff context.

Scoping git log output before scrubbing

Before scrubbing, scope the git log range to the commits you actually need to share. Using git log <base>..<head> limits the output to the relevant range6 rather than the full repository history. A focused log of 30-50 commits is faster to scrub and produces a more compact variables file than a 10,000-commit full history. For release note generation, shrink the scrubbing surface to one release range using the tag range (git log v2.0.0..v2.1.0), limiting scope to exactly the commits included in the release.

Scoping first keeps the variables file compact, because a focused range produces fewer tokens to review. Since CapyToolkit scrubs entirely in your browser, the git history you paste never leaves your machine during processing, and the restored log still contains the real author emails only when you choose to apply the variables file locally.

When to use this

Use this before sharing git log output, git blame results, or git diff text with vendors, AI analysis tools, or external contributors when the history includes developer email addresses or internal references.

Examples

git log with author emails

Before
commit a1b2c3d
Author: [email protected]
Fix connection timeout to db.prod.internal:5432

commit e4f5g6h
Author: [email protected]
Update SMTP config for mailhost.corp
After
commit a1b2c3d
Author: [EMAIL_1]
Fix connection timeout to [DOMAIN_1]:5432

commit e4f5g6h
Author: [EMAIL_2]
Update SMTP config for [DOMAIN_2]

Accidental API key in diff context

Before
-API_KEY=sk-old-production-key-abc123
+API_KEY=sk-new-production-key-def456
After
-API_KEY=[API_1]
+API_KEY=[API_2]

Both the old and new key are tokenized separately, so the diff structure is preserved.

Sources
  1. 1.

    GitHub, "Removing sensitive data from a repository," docs.github.com, accessed June 2026. https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/removing-sensitive-data-from-a-repository

  2. 2.

    Git, "git-diff Documentation," git-scm.com, accessed June 2026. https://git-scm.com/docs/git-diff

  3. 3.

    GitHub, "New options for controlling the default commit message when merging a pull request," github.blog, August 2022. https://github.blog/changelog/2022-08-23-new-options-for-controlling-the-default-commit-message-when-merging-a-pull-request/

  4. 4.

    GitHub, "Creating a commit with multiple authors," docs.github.com, accessed June 2026. https://docs.github.com/en/pull-requests/committing-changes-to-your-project/creating-and-editing-commits/creating-a-commit-with-multiple-authors

  5. 5.

    Git, "git-blame Documentation," git-scm.com, accessed June 2026. https://git-scm.com/docs/git-blame

  6. 6.

    Linux man-pages project, "git-log(1)," man7.org, accessed June 2026. https://man7.org/linux/man-pages/man1/git-log.1.html

FAQ

Yes. Email addresses are one of the 22 detected PII types. Whether in a commit author line, a commit message body, or a code comment, email addresses matching the standard format are tokenized.

Yes. Internal hostnames ending in .corp, .internal, .local, .dev, and .staging are detected. References to internal ticket systems and server names using these TLDs will be tokenized.

The scrubber is a text processor for outbound sharing, not a git history auditor. For systematic credential scanning of your git history, use tools like git-secrets, truffleHog, or GitHub secret scanning. Use the scrubber to clean log output before sharing it.

Paste the git log output into the scrubber, clean it, and then paste the clean version into your AI tool. The AI can still summarize changes, classify commits by type, and generate changelog text from the scrubbed content.

No. Nothing you paste is stored, logged, or transmitted. The tool runs entirely in your browser. Close the tab and all content is gone.

FAQ

You paste it. The scrubber processes text in the browser, so you open the file in an editor or a spreadsheet, copy its contents and paste them into the Scrub tab. Plain-text formats such as CSV, JSON, logs, SQL dumps and git output need no conversion first.

There is no fixed limit, since the work happens in your browser's memory. Very large pastes can make the tab slow, so split big files at record boundaries and keep related records together in each part. Each part then produces its own variables file.

No. It matches the values themselves, so an email address is caught whether it sits in an email column, a nested JSON field, a log message or a commit body. A field name such as customer_email stays in place, because it is not personal data on its own.

Yes, in most cases. Tokens replace values inside existing fields, so delimiters, quotes, brackets and statement syntax stay as they were. A column with a strict type, such as a numeric card number column in a database, may reject a text token, so load scrubbed dumps into text columns or a scratch schema.

No. The variables file maps every token back to its original value, so the two files together rebuild the original data. Send only the scrubbed file, and keep the variables file wherever you keep the source export. CapyToolkit doesn't store either file, so closing the tab without downloading it loses the mapping.

Logs often do, because nobody designed them to hold personal data in the first place. Request lines carry IP addresses, error messages quote user input, and debug logs can include whole request bodies. A CSV export at least shows its personal data in named columns, while a log scatters it through free text.