Scrub PII Before You Paste It Into an AI Assistant
An AI assistant can only leak what it receives. Whatever you paste, upload or leave open in an editor travels to the provider's servers, and from that moment its retention rules, training settings and staff access policies decide what happens to it, not yours. Scrubbing replaces each email address, IP address, API key, card number and connection string with a numbered token such as [EMAIL_1] before the text leaves your browser, so the assistant works on the shape of your problem while the real values stay on your machine.
The workflow is identical for every assistant: paste into the Scrub tab, copy the tokenized text, download the variables file, and paste the assistant's reply into the Restore tab to put the real values back. What changes between assistants is how your text reaches them, and that decides when you scrub. The sections below start with that difference, then cover ChatGPT, Claude, Gemini, GitHub Copilot, Cursor, Grok, Perplexity, Notion AI and code pasted into any of them.
Before you paste into an AI assistant
- Credentials and keys API keys, tokens, connection strings and private key blocks become numbered placeholders
- Contact and network data emails, phone numbers and IP addresses are replaced before the text is copied
- Variables file download it before you close the tab, or the Restore step has nothing to map back
- Names and street addresses pattern matching does not catch them, so replace them by hand
Opens the PII Scrubber with this page's checklist shown at the top of the tool.
Open in the tool →How each kind of assistant receives your text
Scrubbing only works if it happens before the text reaches the assistant, and the three kinds of assistant on this page reach your text at different moments. A chat assistant waits for you to send something, an editor assistant reads files around your cursor on its own, and a workspace assistant reads pages that already sit on the provider's servers. Knowing which kind you use tells you where the scrub step belongs in your day.
Chat assistants receive only what you send
ChatGPT, Claude, Gemini, Grok and Perplexity see the prompt you type, the text you paste and the files you attach, and nothing else from your machine. That makes the scrub step simple: run every paste through the scrubber first, then send the tokenized version. Perplexity adds one wrinkle, because it turns your prompt into web searches, so a value left in the prompt can shape what it goes looking for.
The same rule covers API calls and scripts. If a script builds prompts from logs, tickets or database rows, scrub that text before it goes into the request body, because nothing between your script and the provider will catch a value you forgot. Keep the variables file next to the script's output so you can restore the reply in the Restore tab afterwards.
Editor and workspace assistants read stored text on their own
GitHub Copilot and Cursor collect context from the active file, nearby open files and, when indexing is on, the wider repository, without waiting for a paste. A credential sitting in an open config file can therefore travel with a completion request you never meant to share. On Copilot Business and Enterprise plans, administrators can exclude paths so Copilot ignores those files.1 Cursor reads a .cursorignore file that blocks the listed paths from Agent, Tab and Inline Edit, though its documentation notes that the agent's terminal and MCP tools can still reach them.2
Neither control cleans a file that is already open in an included path, so scrub logs, configs and pasted snippets before they land in the editor, and keep real secrets in environment variables or a secrets manager. Notion AI works one step further back: it reads pages and databases you already store in Notion, so clean meeting notes, HR records and pasted logs before they become pages.
- 1.
GitHub, "Excluding content from GitHub Copilot," docs.github.com, accessed October 2026. https://docs.github.com/en/copilot/how-tos/configure-content-exclusion/exclude-content-from-copilot
- 2.
Cursor, "Ignore File," cursor.com, accessed October 2026. https://prod.cursor.com/docs/reference/ignore-file
Remove Sensitive Data Before Sending to ChatGPT
ChatGPT runs on OpenAI servers. Every prompt you send, including any internal IPs, customer emails, or API keys in your text, is transmitted to and processed by those servers. For developers debugging production code or HR teams drafting sensitive queries, this is a real data governance risk.
By default, OpenAI may use conversations to improve its models unless you opt out or operate under a zero-data-retention agreement.1 Even with retention controls in place, the transmission window exists the moment you click send. Replacing sensitive values with tokens before the prompt leaves your machine closes that window entirely. OpenAI receives only placeholders, never the real data.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →What OpenAI receives from your prompts
ChatGPT processes your full prompt on OpenAI infrastructure. Consequently, internal server addresses, customer email addresses, API keys, and database connection strings travel off your network the moment you submit. Even a casual debugging query such as "why is 10.0.1.42 returning a 503?" discloses an internal IP to a third party. Scrubbing transforms that query to "why is [IP_1] returning a 503?" before it leaves your browser, so OpenAI only ever processes the structural question, not the infrastructure detail.
The transmission happens before any OpenAI retention policy takes effect. Your prompt reaches their servers, gets processed, and a response returns, all within a few seconds. During that window, the raw text of your prompt exists on OpenAI infrastructure. For organizations handling GDPR-regulated data, that transmission alone may constitute a cross-border transfer requiring a legal mechanism under Article 46.2 Scrubbing before the prompt leaves your browser eliminates the identifiable content from that transmission, reducing the data OpenAI receives to structural placeholders that carry no personal or infrastructure information.
Data types the scrubber removes
Around 22 sensitive types are detected automatically: IPv4 and IPv6 addresses, email addresses, all major API key formats (OpenAI sk-, AWS AKIA, GitHub ghp_, Stripe sk_live_/sk_test_, Google AIza, SendGrid SG., Twilio AC, NPM npm_, GitLab glpat-, Slack xox*), JWT tokens, database connection strings, PEM private key blocks, credit card numbers, US Social Security Numbers, phone numbers, and IBAN bank account numbers. Building on this, each unique value receives its own numbered token so the restoration step reconstructs the original exactly.
For ChatGPT specifically, the most commonly exposed types in developer prompts are IPv4 addresses (internal server IPs in debugging queries), API keys (accidentally included in config pastes), and database connection strings (embedded in error messages or stack traces). HR and business teams more frequently expose email addresses, phone numbers, and SSNs when drafting sensitive communications. The scrubber handles all 22 types in a single pass, so you do not need to pre-sort your text by data category before scrubbing.
Restoring real values in ChatGPT responses
After ChatGPT responds with tokens like [IP_1] and [EMAIL_1], paste the response into the Restore tab and upload the variables file you downloaded after scrubbing. The replacement runs locally in under a second. Your team gets a response that references real infrastructure values without OpenAI ever having received them. Conversely, skipping the scrub means the sensitive values are already in OpenAI's systems before any retention policy can act.
The restore step works because the variables file acts as a lookup table. Each token maps to exactly one original value, and the replacement is deterministic: every occurrence of [IP_1] becomes the same IP address it replaced. For long conversations where ChatGPT references multiple tokenized values across several responses, the same variables file handles restoration across all of them. Download the file once after scrubbing, and you can restore any number of ChatGPT responses from that session without re-scrubbing the original text.
ChatGPT tiers and organizational data retention controls
ChatGPT Teams and ChatGPT Enterprise operate under tighter data handling terms than individual accounts. Teams accounts opt out of model training by default for all workspace users, so your organization's conversations are not used to improve OpenAI's models. Enterprise accounts add zero data retention at the request level: OpenAI does not log inputs or outputs after the response is delivered.3 Standard individual Plus accounts operate under broader terms and require manual opt-out under Settings → Data Controls → Improve the model for everyone.
Understanding your account tier matters for how you think about residual risk after scrubbing. Even on Enterprise, the transmission event occurs on every request: your prompt reaches OpenAI's infrastructure, a response is generated, and the session closes. Scrubbing before each prompt ensures that what reaches OpenAI on every transmission contains no identifiable values, making retention controls less critical because there is nothing identifying to retain.
Checking your active tier in the ChatGPT interface
Your subscription tier appears under Settings → Your Plan in the ChatGPT web interface. Teams accounts display "ChatGPT Team" under the plan label. Enterprise accounts require checking the admin console to confirm zero retention is enabled at the organization level. When uncertain about your tier, treat the account as a non-ZDR one and scrub all sensitive values before every prompt regardless.
Custom GPTs and third-party action data flows
Custom GPTs can extend the base ChatGPT interface with external API connections called Actions. When a Custom GPT has Actions enabled, your prompt may trigger a call to a third-party API endpoint chosen by the GPT creator. Any sensitive values in your prompt travel to OpenAI and, if an Action fires, to that third-party endpoint as well.4 Research on plugin and action integrations documents this as a distinct attack surface: because the LLM service composes responses from third-party API output, the boundary between your prompt and an external endpoint is crossed on every tool invocation.4
Scrubbing before interacting with a Custom GPT closes both channels simultaneously. The scrubbed prompt contains no sensitive values, so neither OpenAI's processing nor any triggered Action receives identifiable data. Check the Custom GPT's configuration (visible by clicking the GPT name, then Configure) to see which external services are connected before submitting any prompt that includes internal context or personal data.
Third-party Custom GPTs from the GPT Store
Custom GPTs published in the OpenAI GPT Store can connect to any external service the creator registered. For GPTs built by organizations outside your own, the connected APIs may not be documented publicly. Apply the scrubber for all interactions with third-party Custom GPTs that involve internal or personal data, since the external data flow is less transparent than with first-party OpenAI models.
ChatGPT memory and cross-session persistence of sensitive values
ChatGPT memory stores facts from your conversations and includes them as context in future sessions.5 A sensitive value mentioned once in a debugging session (such as a production hostname or API endpoint) can be saved as a memory and surfaced in completely unrelated future conversations. This mechanism carries data across session boundaries, converting a one-time transmission risk into a persistent context risk.
Disable ChatGPT memory under Settings → Personalization → Memory if your organization restricts sensitive values from persisting in external systems. Alternatively, review and delete specific memories under Settings → Manage Memories after any session that involved sensitive data. Workspace admins on Teams accounts can disable memory organization-wide from the Admin Console.
Scrubbing as a complement to memory management
Memory management and scrubbing address different risks. Memory controls prevent cross-session persistence, while scrubbing prevents single-session transmission of real values. For the highest-risk scenarios involving production credentials or customer records, apply both controls: scrub the prompt before sending and confirm that memory storage is off. The scrubber is the faster of the two steps and requires no account settings to take effect, which makes it the easiest control to apply consistently.
CapyToolkit runs the scrubber entirely in your browser, so this protection applies to every account type without depending on OpenAI retention settings. Because the variables file is the only place the real values live, you keep full control of the data while still getting a complete ChatGPT response after restoration. Applying it takes only seconds before each prompt.
When to use this
Use this before pasting any prompt into ChatGPT that contains internal hostnames, credentials, customer data, or any information covered by your organization's data handling policy.
Examples
DevOps prompt with internal IPs
Why is the connection from 10.0.1.42 to db.corp.internal timing out? The API key is sk-abc123xyz.
Why is the connection from [IP_1] to [DOMAIN_1] timing out? The API key is [API_1].
Paste the scrubbed version to ChatGPT. Download the variables file to restore real values in the response.
HR query with employee data
Draft a performance review for [email protected] who earns $95,000.
Draft a performance review for [EMAIL_1] who earns $95,000.
- 1.
OpenAI, "How your data is used to improve model performance," openai.com, March 2026. https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/
- 2.
Regulation (EU) 2016/679, Article 46, legislation.gov.uk, 2016. https://www.legislation.gov.uk/eur/2016/679/article/46/adopted?view=plain
- 3.
OpenAI, "Enterprise privacy at OpenAI," openai.com, January 2026. https://openai.com/enterprise-privacy/
- 4.
Wanru Zhao, Vidit Khazanchi, Haodi Xing, Xuanli He, Qiongkai Xu, and Nicholas D. Lane, "Attacks on Third-Party APIs of Large Language Models," arXiv:2404.16891, ICLR 2024 Workshop on Secure and Trustworthy Large Language Models, April 2024. https://arxiv.org/abs/2404.16891
- 5.
Kyle Wiggers, "ChatGPT will now use its 'memory' to personalize web searches," TechCrunch, April 2025. https://techcrunch.com/2025/04/18/chatgpt-will-now-use-its-memory-to-personalize-web-searches/
No. The scrubber replaces all detected sensitive values with tokens like [IP_1] before you copy the text. You paste the scrubbed version to ChatGPT. OpenAI only ever sees the tokens, not the original values.
Around 22 types. For developers: IPv4 and IPv6 addresses, API keys (generic sk- prefix, AWS AKIA, GitHub ghp_, GitLab glpat-, Stripe sk_live_/sk_test_, Google AIza, SendGrid SG., Twilio AC, NPM npm_, Slack xox*), JWT tokens, database connection strings, and PEM private key blocks. For business and HR use: email addresses, credit card numbers, US Social Security Numbers, phone numbers, IBAN bank account numbers, and internal domain names.
Yes. Download the variables file after scrubbing, then paste ChatGPT's response into the Restore tab and upload the file. Every token is replaced with the original value.
No. CapyToolkit is an independent browser-based utility with no connection to OpenAI. The scrubber runs entirely in your browser, processes your prompts locally, and never interacts with OpenAI or any other AI provider.
No. All processing runs locally in your browser using JavaScript. No text is transmitted to any server. Verify this by opening DevTools and watching the Network tab while you paste.
Remove Sensitive Data Before Sending to Claude
Every prompt you send to Claude leaves your network. Internal credentials, customer emails, production IP addresses: all of it arrives on Anthropic's servers, where their data retention and model training policies govern what happens next.
Enterprise agreements and zero-data-retention options exist, but most developers and small teams using Claude.ai or the standard API operate without a negotiated policy. Scrubbing sensitive values before the prompt reaches Claude's servers eliminates the transmission risk regardless of which tier you use.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →How Claude uses submitted data
By default, Anthropic may use API and Claude.ai conversations to improve its models unless you opt out through a business agreement.1 Consequently, a prompt containing your production database URL or an employee's Social Security Number could become training material for a future model. Even with zero-data-retention set, the transmission itself creates a window of exposure that smart pre-processing eliminates entirely. Scrubbing sensitive values before the request fires means the risk window never opens.
The transmission event is the critical moment. Your prompt leaves your network, reaches Anthropic's infrastructure, and gets processed by their models. Every value in that prompt exists on their servers during processing, regardless of how briefly. For organizations subject to GDPR, CCPA, or internal data governance policies, that transmission window is the exposure point.2 Scrubbing before the prompt travels across the network ensures that only tokens reach Anthropic's infrastructure, so the processing windows contains no identifiable data and the retention policy question becomes irrelevant.
What Claude needs versus what you send
Claude understands context, not raw values. When you write "debug why 10.0.1.42 returns a 503," you could write "debug why [IP_1] returns a 503" and receive exactly the same quality of analysis. Building on this, token substitution preserves the semantic structure of every prompt, including the relationship between entities, the shape of the problem, and the relevant code path, while removing the identifiable values that create compliance exposure. Claude never needed the real IP.
This principle applies universally across AI models: the structural content of a prompt drives the quality of the response, not the specific values. An AI that sees "[API_1] is used to authenticate requests to [DOMAIN_1]" can still identify that hardcoded credentials in source code are a security risk and suggest using environment variables instead. The real API key and domain name add no analytical value. Removing them before the prompt reaches Claude preserves your infrastructure confidentiality without reducing the quality of the assistance you receive.
Restoring real values in Claude responses
After Claude responds with [IP_1] in its output, paste the response into the Restore tab and upload your downloaded variables file. The scrubber replaces every placeholder with the original value in under a second. Your team gets a response that reads naturally and contains your real server addresses, without ever having exposed them to an external server. Consequently, the two-way workflow lets Claude function as a fully capable debugging partner while your data governance policy remains intact.
The restore workflow is especially useful for teams that interact with Claude across multiple turns in a single conversation. Each turn from Claude may reference tokenized values from your original prompt. Rather than downloading the variables file after every message, download it once after your initial scrub and keep it available for restoring any response in the conversation. When the conversation is complete, the variables file serves as a record of what was tokenized, which is useful for audit trails in organizations that document data sharing with AI providers.
Anthropic data retention tiers: consumer, API, and enterprise
Anthropic offers three access tiers with distinct data handling characteristics. Claude.ai consumer accounts operate under Anthropic's standard privacy policy, which permits using conversations to improve models unless you opt out under Account Settings → Privacy → Improve Claude for everyone. Standard API accounts without a negotiated data agreement operate under API terms that also permit training use by default. Claude Enterprise plans include a zero-data-retention (ZDR) option that prevents Anthropic from storing conversation content after the response is delivered.3
Understanding which tier applies to your organization changes how you think about residual risk after scrubbing. Under ZDR, the transmission window is the only exposure: your prompt reaches Anthropic's infrastructure, the response is generated, and nothing persists. Even with ZDR active, the transmission event occurs on every request. Scrubbing ensures that what transmits contains no identifiable values, making the retention question irrelevant for the fields you have removed.
Verifying ZDR status for Claude Enterprise accounts
ZDR must be explicitly configured at the organization level in the Anthropic Console.3 Individual developers using a personal API key do not have ZDR access regardless of their billing tier. If you are uncertain which tier applies to your API key, check the organization settings in Anthropic Console under Security → Data retention. When uncertain, treat the account as non-ZDR and scrub all sensitive values before every request.
Multi-turn conversations and context accumulation risk
Claude.ai maintains conversation history within each session. Every turn you add to a conversation appends to the context window transmitted to Anthropic's servers on subsequent turns. Consequently, a sensitive value included in turn 3 of a conversation is included in the full context sent on turns 4, 5, and every turn that follows. Scrubbing each turn before you send it prevents sensitive data from accumulating in the conversation thread.
The risk compounds in long debugging sessions where each turn builds on prior responses. A developer who includes a real database connection string in turn 1 exposes that value in every subsequent request in the session, even if later turns contain no new sensitive data. Starting a fresh conversation for each topic, or scrubbing at the first mention of any sensitive value, prevents this compounding effect entirely.
Saved conversation history and post-session cleanup
Claude.ai saves conversation history in your account.4 Any sensitive value sent without scrubbing in a past conversation persists in that history until you delete the specific conversation. For teams with strict data governance requirements, establish a policy of deleting Claude.ai conversations that involved internal data after the session closes. The more reliable long-term approach is to make scrubbing the habit, not conversation cleanup, so the history contains only tokens rather than real values and no post-session cleanup is needed.
System prompts and operator-level data exposure in the Claude API
Organizations using the Claude API directly build system prompts that establish context, persona, and instructions for every user interaction. System prompts frequently include internal tool descriptions, internal API endpoint URLs, and example data with real infrastructure references embedded. These system prompts are part of every API request and transmit to Anthropic's servers on each call.5
Apply the scrubber to system prompt content before it enters production. A system prompt that references your internal CRM API endpoint as context sends that endpoint to Anthropic on every single API call your application makes. Tokenizing internal references in the system prompt and maintaining a separate mapping for operator context removes this exposure from every user interaction automatically, rather than requiring prompt-level scrubbing at runtime.
Credential management for multi-tenant Claude API applications
Multi-tenant applications that include organization-specific context in system prompts create a more complex exposure pattern. Each tenant's system prompt may include their specific API endpoints, internal domain names, or example data. Scrubbing each tenant's system prompt template at deployment time, rather than at runtime, gives you a stable and auditable baseline: the variables file for each tenant's prompt documents what was replaced, creating a record suitable for security review. Store per-tenant variables files in your secrets management system alongside the original credentials they map to.
Treating each tenant's template as a scrubbed artifact means the deployed system prompt carries no live credentials into Anthropic's infrastructure. CapyToolkit runs this step in your browser, so the scrubbing happens before any request leaves the machine and the variables file stays under your own access controls rather than in a shared build environment.
When to use this
Use this before pasting any prompt into Claude that contains production credentials, customer data, internal hostnames, or content covered by your team's data handling or AI usage policy.
Examples
Production error with internal hostname
Our service at auth.internal.company.com returns 401 when the JWT eyJhbGciOiJSUzI1NiJ9.eyJ1c2VyX2lkIjoiMTIzIn0.sig is passed.
Our service at [DOMAIN_1] returns 401 when the JWT [JWT_1] is passed.
Paste the scrubbed version into Claude. The analysis is identical, and Claude never needed the real hostname or token.
Database query with credentials in the connection string
Explain why this query is slow: SELECT * FROM users WHERE email = '[email protected]' — db is at postgres://admin:[email protected]/main
Explain why this query is slow: SELECT * FROM users WHERE email = '[EMAIL_1]' — db is at [DBURL_1]
- 1.
Anthropic, "Is my data used for model training?," privacy.claude.com, March 2026. https://privacy.claude.com/en/articles/10023580-is-my-data-used-for-model-training
- 2.
European Data Protection Board, "Report of the work undertaken by the ChatGPT Taskforce," edpb.europa.eu, May 2024. https://www.edpb.europa.eu/documents/task-force-report/report-of-the-work-undertaken-by-the-chatgpt-taskforce_en
- 3.
Anthropic, "Zero Data Retention," code.claude.com, accessed June 2026. https://code.claude.com/docs/en/zero-data-retention
- 4.
Anthropic, "How long do you store my data?," privacy.claude.com, July 2026. https://privacy.claude.com/en/articles/10023548-how-long-do-you-store-my-data
- 5.
OWASP Foundation, "LLM02:2025 Sensitive Information Disclosure," owasp.org, 2025. https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/main/2_0_vulns/LLM02_SensitiveInformationDisclosure.md
Anthropic's default policy allows use of Claude.ai conversations for model improvement unless you opt out. API users with a business agreement can request zero data retention. Regardless of policy, scrubbing sensitive values before sending ensures nothing identifiable ever reaches Anthropic's servers.
Yes. Claude understands that [IP_1] is a placeholder for a real address. It can diagnose connection issues, suggest query optimizations, and review security logic without ever knowing the actual values.
Around 22 types: IPv4 and IPv6 addresses, API keys, AWS access keys, GitHub and GitLab tokens, Stripe and Google API keys, SendGrid and Twilio keys, NPM tokens, JWT tokens, database connection strings, PEM private key blocks, email addresses, credit card numbers, SSNs, phone numbers, and IBAN bank account numbers.
No. All processing runs locally in your browser using JavaScript. No text leaves your device. You can verify this by opening DevTools and watching the Network tab while you paste.
Yes. CapyToolkit requires no sign-up or login to scrub your prompts. Open the tool, paste your text to remove PII before sending to Claude, copy the cleaned output, and scrub again for the next task. Nothing is stored between sessions.
Remove Sensitive Data Before Sending to Google Gemini
A single prompt to Gemini can carry more sensitive data than you realize. Internal IP addresses in debugging queries, customer emails in support drafts, API keys in config examples: every value in your prompt travels to Google's servers the moment you hit send.
Workspace accounts benefit from Google's data processing controls, but consumer and standard API accounts operate under broader terms.1 Regardless of your account tier, the scrubber ensures sensitive values never reach Google's infrastructure. Tokenizing before submission closes the transmission window at the point of input.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →How Gemini handles submitted data
Gemini processes your full prompt on Google infrastructure. For Google One AI Premium and Workspace plans, retention settings reduce how long conversations are stored, but the transmission still occurs. Consequently, a prompt containing your production database URL or a customer's Social Security Number reaches Google's servers under their current data terms, and the sender, not Google, bears responsibility for that transmission under GDPR and CCPA.2 Removing the sensitive values before the prompt eliminates this responsibility exposure regardless of your account tier.
Google's data processing terms for Workspace include specific provisions around AI data handling: Workspace data is not used to train Google's general-purpose AI models.1 However, this protection applies at the account configuration level and requires the organization's admin to have enabled the relevant data protection settings. For consumer Google accounts and standard API access, the data handling terms are broader. Scrubbing before submission provides protection that does not depend on which Google account type you use or which admin settings are active.
Network and security data in Gemini prompts
IT and security teams frequently use Gemini to analyze network configurations, security policies, and infrastructure architecture. These queries often include internal IP addresses, internal domain names (.corp, .internal, .staging), and API credentials embedded in config examples. Building on this, developers paste stack traces and error logs that contain connection strings and authentication tokens. The scrubber detects all of these formats and tokenizes the values that identify your specific infrastructure before they reach Google.
Security teams face a particular challenge: the queries most useful for AI-assisted analysis are the ones most likely to contain sensitive infrastructure details. A prompt asking Gemini to review a firewall rule set for misconfigurations needs the actual rule text to be useful, and that rule text contains internal IP ranges and service hostnames. Scrubbing lets you submit the full structural content for analysis while replacing only the identifying values. Gemini receives the rule logic, port numbers, and protocol details it needs for a thorough review, and your internal addressing scheme stays private.
Using the variables file with Gemini responses
After Gemini responds with [IP_1] or [DOMAIN_1] in its output, switch to the Restore tab, paste the response, and upload the variables file downloaded after the initial scrub. The restoration replaces every token with its original value in your browser. Your team reads a natural-language response referencing real infrastructure. Yet the variables file itself should be treated as sensitive because it is the key that reconstructs the original data, so access to it should be restricted to the same people who could access the original prompt.
For teams running multiple Gemini sessions in a day, establish a routine of downloading the variables file immediately after each scrub and storing it in a location with the same access controls as the original data. The variables file is small, typically a few kilobytes, and can be attached to the same ticket or document where you track the AI-assisted analysis. When the analysis is complete and the response has been restored, delete the variables file if your data retention policy requires it. Treating the variables file as a transient key rather than a permanent record keeps your data lifecycle clean.
Gemini for Google Workspace versus consumer gemini.google.com
Gemini's data handling differs significantly between personal Google accounts and Google Workspace accounts. Consumer accounts at gemini.google.com operate under Google's standard privacy policy, which includes conversation review for safety and quality improvement.3 Workspace accounts on Business, Enterprise, or Frontline plans use Gemini for Google Workspace, which operates under the Google Workspace Data Processing Amendment: Google does not use your organization's data to train foundational AI models, and interactions are covered by the Workspace DPA.
For teams on Google Workspace, the relevant admin controls live under Admin console → Apps → Additional Google Services → Gemini. Workspace admins can restrict which organizational units access Gemini, limit which AI features are enabled, and control conversation export settings. Regardless of these admin choices, prompts are still transmitted to Google's servers for processing. Scrubbing before submission remains the correct approach for sensitive data, because the transmission event occurs regardless of which admin policies are active.
Confirming your Gemini account context before sharing internal data
Before relying on Workspace data protections, confirm you are accessing Gemini through your Workspace account rather than a personal Gmail account that shares an email domain with your organization. Signing into gemini.google.com with a personal Google account does not apply Workspace DPA protections even if the email domain matches. Check the account icon in the Gemini interface to confirm the active account is your Workspace identity, and when uncertain, scrub all sensitive fields before submission regardless of account context.
File and document upload exposure in Gemini
Gemini accepts file uploads alongside text prompts: PDFs, spreadsheets, images, and audio files.4 When you upload a document for summarization or analysis, the full document content is transmitted to Google's servers as part of the request context. A PDF containing patient records, a spreadsheet with employee salary data, or a network diagram with labeled internal IP addresses all travel to Google's infrastructure when uploaded.
For document-level submissions, extract PDF text before scrubbing, not the file. PDF text can be copied with your system clipboard or exported using a local PDF reader. After scrubbing the extracted text, paste the cleaned version directly into the Gemini prompt field rather than uploading the original file. This approach gives Gemini the structural content it needs for analysis without sending the original document with all its embedded identifiers.
Redacting visual content in screenshots and diagrams
Screenshots and infrastructure diagrams often contain visible IP addresses, hostname labels, and credentials as plain text within the image. Before uploading any image to Gemini that shows internal infrastructure, review the visual content for text labels. Use an image editor to blur or cover sensitive labels before upload, or export a version of the diagram with placeholder names replacing real hostnames and IPs rather than uploading the original.
Gemini integrations in Google Docs, Gmail, and Workspace apps
Gemini features embedded in Google Docs, Sheets, and Gmail send the surrounding document or email content as context alongside your explicit instruction. When you invoke "Help me write" or "Summarize" in a Docs document that contains meeting notes with attendee names and internal server references, that full page context travels to Gemini for processing. The AI feature sees more than your typed instruction; it sees the surrounding content in the document or email thread.5 Google documents this behavior directly: when editing in Docs, Gemini understands the context of the document even when no text is selected.5
Limiting context when invoking Gemini on sensitive documents
For documents that contain PII, use a targeted prompt approach rather than full-context Gemini features. Instead of invoking Gemini directly on a sensitive document, copy only the structural elements (headings, section outlines, non-sensitive context) into a fresh document or paste them into gemini.google.com after scrubbing. This limits the context Gemini receives to what is structurally necessary for your task without exposing the PII-heavy content from the original document.
Because the scrubber processes text locally in your browser, this targeted approach works the same way for any Google account, from a personal Gmail login to a Workspace Enterprise plan. The structural copy you paste into gemini.google.com keeps the document useful for Gemini while the variables file holds the real names and identifiers for later restoration.
When to use this
Use this before pasting any prompt into Gemini that contains data your organization classifies as internal, confidential, or personally identifiable.
Examples
Network debugging prompt
Explain why traffic from 192.168.10.5 to api.internal-corp.com port 443 returns a 503.
Explain why traffic from [IP_1] to [DOMAIN_1] port 443 returns a 503.
Code review prompt with credentials
Review this config: host=db.prod.internal user=admin password=s3cr3t!
Review this config: host=[DOMAIN_1] user=admin password=[API_1]
- 1.
Google Cloud, "Service Specific Terms," Section 18, cloud.google.com, accessed June 2026. https://cloud.google.com/terms/service-terms
- 2.
Regulation (EU) 2016/679, Article 46, legislation.gov.uk, 2016. https://www.legislation.gov.uk/eur/2016/679/article/46/adopted?view=plain
- 3.
Kyle Wiggers, "Google saves your conversations with Gemini for years by default," TechCrunch, February 2024. https://techcrunch.com/2024/02/08/google-saves-your-conversations-with-gemini-for-years-by-default/
- 4.
Google, "Upload & analyze files in Gemini Apps," support.google.com, accessed June 2026. https://support.google.com/gemini/answer/14903178
- 5.
Google, "Write & edit with Gemini in Docs," support.google.com, accessed October 2026. https://support.google.com/docs/answer/13447609
Google's data retention policies vary by account type and workspace settings. Regardless of their policy, if the sensitive data never reaches their servers, there is nothing for them to retain. The scrubber ensures only tokens reach Gemini.
Yes, as long as you run the same source text through the scrubber each session. The tokens are deterministic for the same input text. However, to be safe, regenerate the variables file for each sensitive session.
Yes. The scrubber works with any AI tool because it is a browser-side text preprocessor. Copy the scrubbed output and paste it into any AI interface.
No. The scrubber does not collect, store, or transmit any text you paste. All processing runs locally in your browser.
No. CapyToolkit is an independent browser-based utility with no connection to Google.
Remove Sensitive Data Before Using GitHub Copilot
Unlike most AI tools, GitHub Copilot does not wait for you to paste anything. It reads your active file, adjacent files, and recent edits automatically, bundling that context into every request it sends to Microsoft's servers. A config file open in a background tab, a commented-out credential in a test fixture, a .env example in your project tree: Copilot sees all of it.
Content exclusion policies in Copilot Business and Enterprise can block specific file paths, but the default free and individual tiers send context broadly.1 For developers working with production configs or internal infrastructure definitions, a pre-paste scrub removes sensitive elements before they enter the editor context that Copilot reads.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →What Copilot reads from your editor
Copilot draws context from the currently open file, adjacent files in the same project, and recently edited files in the session.2 Consequently, pasting a config file with a database password into any file in your open project, even a scratch file, creates a window where that credential exists in Copilot's context. Yet the developer often isn't aware of what context Copilot is bundling, because the process is automatic and invisible in the sidebar. Scrubbing credentials before they enter the editor is the only reliable way to keep them out of that context.
The context window Copilot uses is larger than most developers expect. In VS Code, Copilot can reference files that are not currently visible in any editor tab, including files that were opened earlier in the session and then closed. A developer who opens a .env file to check a variable name, switches to a source file, and then asks Copilot for help may inadvertently send the .env file contents as part of the context. Scrubbing before pasting into any file in the project eliminates this passive exposure path entirely.
Detecting credentials in developer context
The scrubber detects all major credential formats that appear in developer files: database connection strings (postgres://, mysql://, mongodb://, redis://), AWS access keys (AKIA prefix), GitHub PATs (ghp_), GitLab tokens (glpat-), Stripe payment keys (sk_live_, sk_test_), and generic API keys (sk- prefix). Building on this, it also catches JWT tokens, PEM private key blocks, and internal hostnames (.corp, .internal, .local, .staging). These are exactly the formats that appear in .env files, docker-compose.yml, deployment scripts, and CI/CD configs, which are the files developers most commonly edit with Copilot active.
The detection covers both active credentials and commented-out examples. Developers frequently leave old credentials in comments as documentation of what the config used to contain. These commented-out values are still valid credential patterns and are still detected by the scrubber. Before pasting a config file into your editor for Copilot-assisted editing, run it through the scrubber to catch both active and historical credentials. This is especially important for files with long comment histories that document configuration changes over time.
Enterprise policy versus technical control
GitHub Copilot Business and Enterprise offer content exclusion policies that prevent specific files or patterns from being sent. However, configuring exclusion policies requires admin access and careful policy maintenance, and a gap opens every time a new pattern is introduced. Conversely, running text through a pre-paste scrubber requires zero configuration, works in every editor, and applies consistently regardless of Copilot subscription tier. Technical controls that run before the data enters the pipeline are more reliable than policies that filter after the fact.
Content exclusion policies and pre-paste scrubbing address different parts of the problem. Exclusion policies prevent Copilot from reading specific files as passive context, but they do not prevent a developer from manually copying content from an excluded file and pasting it into a prompt. The scrubber addresses the active paste path: even if a developer intentionally copies content from a secrets file, scrubbing that content before it enters the editor ensures Copilot never sees the real values. Using both controls together provides defense in depth against both passive context capture and active paste-based exposure.
Content exclusion policies in GitHub Copilot for Business
GitHub Copilot for Business allows workspace admins to define content exclusion rules that prevent specific files and directories from being included in Copilot's context window.3 You configure exclusions in the repository settings under Settings → Copilot → Content Exclusions, using glob patterns to match file paths. A pattern like **/.env* excludes all .env files across the repository, preventing their contents from being sent upstream even when those files are open in an active editor session.
Content exclusion rules apply at the repository level and require the repository to be covered by a Copilot for Business or Copilot Enterprise subscription managed by an organization admin. Individual Copilot for Individuals subscribers cannot configure content exclusions. Maintaining accurate exclusion rules demands ongoing attention: a new credential file added to the repository with a non-standard name (for example, secrets.yaml instead of .env) remains uncovered by existing patterns until an admin manually adds the new pattern.
Closing the gap between content exclusion policies and paste-based exposure
Content exclusions act on file paths, not on content. A secrets file with an unexpected path bypasses all exclusions. Additionally, exclusions prevent Copilot from sending the file as passive context, but they do not prevent a developer from manually copying content from that file and pasting it into the Copilot Chat input field. Running a pre-paste scrub closes this second path. Combining content exclusion policies with a scrub-before-paste discipline provides defense in depth: the policy blocks passive context capture, while the scrubber blocks active paste-based exposure.
Copilot Chat versus inline completions: different context scopes
Copilot's inline completion feature draws context from the immediately surrounding code in the active file and from recently edited files in the session. Copilot Chat expands this scope significantly: in VS Code, Copilot Chat can reference the entire workspace using the @workspace participant, which indexes all files in the open folder for retrieval.4 Chat messages that include @workspace trigger a broader context fetch that may include files not currently open in any editor tab.
The #file and #selection references in Copilot Chat let you explicitly add specific files or code selections to a prompt. Using #file:secrets.yaml directly attaches that file's content to the next prompt. Developers who explicitly reference credential files this way bypass content exclusion policies, because exclusion rules filter passive context collection rather than explicit user-initiated file references, which is exactly why you should scrub before using #file on a secrets file rather than trust policy alone.
How Copilot Workspace extends context across entire repositories
GitHub Copilot Workspace allows you to open a GitHub issue and have Copilot plan and implement code changes across the whole repository. The entire repo tree is indexed for planning, including any files not covered by your editor's content exclusion configuration.4 Before using Copilot Workspace on a repository that contains credential files, confirm that those files are listed in your .gitignore and have never been committed, since Copilot Workspace indexes the committed file tree rather than just what is currently open in your editor.
Building a scrub-before-paste habit for Copilot workflows
Establishing a pre-paste scrub as a routine habit is the simplest governance control for Copilot users. Your scrub step takes under 30 seconds: paste the text into the browser scrubber, copy the scrubbed output, then paste that into your editor or Copilot Chat. Repeating this before every paste from an external source (a config file, a log file, a support ticket) eliminates the entire category of accidental credential disclosure to Copilot without requiring any policy configuration, admin access, or subscription upgrade.
Team adoption improves when the scrubber URL is included directly in your team's AI tool usage guide or internal wiki page documenting AI governance. Linking to the scrubber in the same location where you document Copilot configuration ensures developers who set up Copilot also discover the pre-paste step at the same time. The scrubber is free and requires no account, which removes friction to adoption from day one.
Monitoring Copilot context for credential exposure signals
GitHub Copilot for Business provides usage telemetry through the GitHub Copilot usage APIs, which return aggregated and user-specific reports covering feature usage, engagement, and adoption.5 The returned reports carry counts and rates for suggestion generation and acceptance, which give you acceptance rates and active seat counts but no access to the content of suggestions or context windows. For teams that want signal on whether credentials are reaching Copilot, a browser DLP extension that intercepts HTTPS requests to copilot-proxy.githubusercontent.com and scans request payloads for credential patterns provides the most direct approach to post-hoc monitoring.
Pairing this monitoring with the browser scrubber closes the loop, because the scrubber stops the credential from entering Copilot's context in the first place. Since CapyToolkit processes text locally with no server upload, the detection runs before the request is built, giving you a preventive control that complements the post-hoc telemetry the endpoint reports.
When to use this
Use this before pasting any config file, .env example, log output, or infrastructure definition into an editor where GitHub Copilot is active, especially if the file contains credentials or internal hostnames.
Examples
.env file with live credentials
DATABASE_URL=postgres://admin:[email protected]:5432/app STRIPE_KEY=sk_live_xyz123 AWS_KEY=AKIAIOSFODNN7EXAMPLE
DATABASE_URL=[DBURL_1] STRIPE_KEY=[STRIPE_1] AWS_KEY=[AWS_1]
Paste the scrubbed version into your editor. Copilot sees only the variable names and tokens.
docker-compose.yml with internal service addresses
environment: REDIS_URL: redis://redis.internal.corp:6379 API_KEY: sk-prod-service-key-abc123
environment: REDIS_URL: [DBURL_1] API_KEY: [API_1]
- 1.
GitHub, "Excluding content from GitHub Copilot," docs.github.com, accessed June 2026. https://docs.github.com/en/copilot/how-tos/configure-content-exclusion/exclude-content-from-copilot
- 2.
Microsoft, "Manage chat context in GitHub Copilot Chat," learn.microsoft.com, accessed June 2026. https://learn.microsoft.com/en-us/visualstudio/ide/copilot-chat-context-references?view=vs-2022
- 3.
GitHub, "GitHub Copilot Workspace: Welcome to the Copilot-native developer environment," github.blog, April 2024. https://github.blog/news-insights/product-news/github-copilot-workspace/
- 4.
GitHub, "Indexing repositories for GitHub Copilot," docs.github.com, accessed June 2026. https://docs.github.com/en/copilot/concepts/context/repository-indexing
- 5.
GitHub, "Track organization Copilot usage," github.blog, December 2025. https://github.blog/changelog/2025-12-16-track-organization-copilot-usage/
Copilot sends a subset of your current editor context, including the active file and nearby files, to GitHub's servers to generate suggestions. The exact scope depends on your subscription and workspace settings. To be safe, assume any file open in your editor can contribute to Copilot's context.
Content exclusion policies reduce risk but require correct configuration and ongoing maintenance. A pre-paste scrub is an additional, complementary control that works regardless of your subscription tier and without relying on policy updates.
Yes. Paste the entire file into the scrubber. Every credential format is detected in one pass. Download the variables file, then paste the scrubbed version into your editor.
Yes. The scrubber is a browser-based text processor. It works the same for VS Code, JetBrains, Vim, or any other editor. Paste your text into the browser, copy the scrubbed output, then paste it into your editor.
No. The scrubber runs as client-side JavaScript in your browser tab. No text is transmitted to any server. You can verify this by opening DevTools, going to the Network tab, and confirming zero outbound requests while you paste.
Remove Sensitive Data Before Using Cursor
Cursor sends code context to frontier AI models by design. The editor bundles your current file, recent edits, terminal output, and optionally entire codebase indexes into prompts it sends to models including Claude and GPT.1 Developers using Cursor on production systems must assume that anything in their editor can reach these providers.
Unlike traditional editors, Cursor's AI features actively read your project to improve suggestion quality. Consequently, pasting a secrets-laden config file, a .env example, or log output containing internal IPs into a Cursor project creates a direct path for those credentials to reach AI providers. Pre-scrubbing before pasting into Cursor removes this risk at the source.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →How Cursor shares context with AI models
Cursor's Chat and Composer features send your current file, selected code, and optionally additional files to the configured AI provider. The Codebase Indexing feature ingests your entire repo into an embedding store, which Cursor queries to build relevant context for each AI request.2 Building on this: if your repo contains hardcoded credentials in any file, even a commented-out example or a test fixture, those values can be included in the context sent upstream. The scrubber eliminates the risk by removing credentials before they enter the project.
Codebase indexing is the feature that makes Cursor powerful and risky simultaneously. When you index a repository, every file becomes searchable context for AI queries. A credentials file committed months ago, even if removed in a later commit, may still exist in the index until you rebuild it. Before indexing a repository that has ever contained credentials, scrub the sensitive files first or add them to .cursorignore. The scrubber handles the text-level removal; .cursorignore handles the index-level exclusion. Together they close both paths.
What the scrubber detects in developer files
For developer content, the scrubber detects all major credential formats: AWS access keys, GitHub and GitLab PATs, Stripe payment keys, SendGrid and Twilio keys, NPM auth tokens, generic API keys, JWT tokens, PEM private key blocks, database connection strings, and internal hostnames. Furthermore, it catches IPv4 and IPv6 addresses and email addresses, which are common in stack traces, CI logs, and error messages that developers paste into Cursor for debugging help. Each unique value gets its own numbered token so the restoration step is exact.
Cursor users frequently paste terminal output, CI logs, and error traces into Chat for debugging assistance. These text sources are rich in sensitive data: a single failed deployment log can contain database connection strings, internal IP addresses, and authentication tokens all in one paste. The scrubber processes this unstructured log output the same way it processes code: paste the log text, and every detected credential pattern gets tokenized. The AI receives the error structure and diagnostic context it needs for debugging without receiving any of your production secrets.
Privacy mode and its limits
Cursor offers a Privacy Mode setting that promises not to store code on Cursor's servers or use it for training.1 Yet even with Privacy Mode enabled, your code is still transmitted to the underlying AI provider (Claude, GPT-4, etc.) during the request. Privacy Mode controls what Cursor does with the data, not what the AI provider does. Independent scoring of major coding assistants finds that none of them filter secrets out of user prompts before transmission, leaving that responsibility entirely with the developer.3 Consequently, organization-level AI usage policies that restrict sending production credentials to any external service apply even in Privacy Mode. Scrubbing before pasting satisfies those policies regardless of Privacy Mode status.
Think of Privacy Mode as a control over Cursor's behavior, not the AI provider's behavior. Your code reaches the AI model provider regardless of Cursor's privacy settings, because the model provider needs the code context to generate a response. If your organization's policy restricts sending production credentials to any external service, the policy applies to the AI model provider as much as it applies to Cursor. Scrubbing before pasting is the only control that enforces the policy at the data level, independent of which Cursor settings are configured.
Cursor's codebase indexing: embeddings, models, and data destinations
Cursor's Codebase Indexing feature generates vector embeddings of your source files using an embedding model served through Cursor's API infrastructure.2 Cursor sends file contents to its embedding service, which returns vector representations stored locally in .cursor/index/. The indexed embeddings let Cursor perform similarity search across your codebase when assembling context for Chat and Composer requests. Files indexed at any point remain retrievable for context assembly until you delete or rebuild the index.
You control which files enter the index using a .cursorignore file at the project root, which follows the same glob syntax as .gitignore.4 Adding *.env, secrets.yaml, and *.pem to .cursorignore prevents those files from being read by Agent, Tab, Inline Edit, or @ mention references. Cursor also ignores .gitignore patterns and .env files by default, and its own documentation recommends patterns such as **/.env, **/credentials.json, and **/*.pem specifically to restrict access to API keys and credentials. Two gaps remain. The terminal and MCP tools run outside Cursor's file access controls, so an ignored file can still be read by a shell command or a connected MCP server. And Cursor notes that while ignored files are blocked, complete protection is not guaranteed because of LLM unpredictability.
Keeping credential files out of the Cursor index permanently
Place credential files outside the project root entirely, or use a secrets management tool such as HashiCorp Vault, AWS Secrets Manager, or the 1Password Secrets Automation SDK to inject credentials at runtime rather than storing them in files within the project directory. When credential files must exist in the project directory, add their exact paths to .cursorignore before running Cursor for the first time on that project. Test your patterns with git check-ignore -v [file], which uses the same matching engine Cursor documents for ignore files.
AI model selection in Cursor and its effect on data routing
Cursor routes inference requests to the selected AI model. Each model option sends your context to a different provider: Claude models route to Anthropic's API, GPT-4 and o3 models route to OpenAI's API, and Gemini models route to Google's API.5 Switching models in the top-right model picker changes the data destination for every subsequent request without a separate notification about the privacy policy change. A developer who uses Claude for most sessions and switches to GPT-4 for a specific task sends that session's context to OpenAI rather than Anthropic.
Your organization's AI governance policy may approve different AI providers for different data categories. Code analysis using open-source algorithms may be permissible with a broader set of providers than analysis involving proprietary business logic or internal service credentials. Scrubbing credentials and internal hostnames before pasting into Cursor Chat eliminates the provider-specific routing concern: the context that reaches any provider contains no identifiable sensitive values regardless of which model is selected.
Cursor Business plan and organizational data controls
Cursor Business accounts provide a centralized admin panel where admins can enforce model availability, disable personal API key use, and enable privacy mode across the organization. Business plan privacy mode promises that code is not retained by Cursor for training and is not sent to any third party beyond the AI model provider. Verify these settings at cursor.sh/settings before assuming business plan protections are active for your organization's seats, as they require explicit enablement rather than being on by default.
Terminal integration and diff view: additional data inputs to Cursor AI
Cursor integrates an AI-accessible terminal that lets you run shell commands and reference the output in Chat messages using the @Terminal reference. When you run a command in the Cursor terminal and then type @Terminal in Chat, the terminal output is attached to the prompt and sent to the active AI model. Terminal output from deployment scripts, database migrations, and CI runner commands frequently contains internal hostnames, connection strings, and authentication tokens embedded in error messages.
The diff view in Cursor, accessible through the Source Control panel, shows file changes alongside an AI Chat interface. Requesting AI review of a diff that includes credential changes (for example, rotating an API key in a config file) sends both the old and new credential values as diff context to the AI provider, so clean a credential-rotation diff before AI review to keep neither value from ever reaching the model.
Pasting CI/CD output into Cursor for debugging
CI/CD pipeline output accessed via GitHub Actions run logs, GitLab CI job output, or Buildkite build artifacts frequently contains environment variable values in error messages, command traces produced by set -x, and initialization scripts that echo configuration. Before pasting any CI output into a Cursor terminal or Chat message, run it through the scrubber. Pipelines using trace mode are particularly risky because every environment variable name and value appears in the debug output, making pre-paste scrubbing an essential step for any developer using Cursor for CI failure analysis.
Running every CI log through the scrubber before it reaches the Cursor terminal removes the highest-risk tokens at the source. Because CapyToolkit works entirely in your browser, the paste never leaves your machine during scrubbing, and the variables file you download lets you restore the real values locally once Cursor has returned its debugging analysis.
When to use this
Use this before pasting any file, log, config, or code snippet into a Cursor project that you want to keep isolated from external AI providers, especially production credentials, internal hostnames, or customer data.
Examples
Log output with credentials pasted for debugging
ConnectionError connecting to redis://cache.prod.internal:6379 with password abc123xyz — from IP 10.20.30.40
ConnectionError connecting to [DBURL_1] with password [API_1] — from IP [IP_1]
Paste the scrubbed log into Cursor Chat. You get the same debugging analysis without the production endpoint reaching the AI.
Settings file with payment and external service keys
STRIPE_LIVE_KEY=sk_live_51Ab SENDGRID_KEY=SG.xyz123 INTERNAL_API=sk-internal-service-key
STRIPE_LIVE_KEY=[STRIPE_1] SENDGRID_KEY=[SENDGRID_1] INTERNAL_API=[API_1]
- 1.
Cursor, "Privacy and data," cursor.com, accessed June 2026. https://cursor.com/help/security-and-privacy/privacy
- 2.
Cursor, "Securely indexing large codebases," cursor.com, 2025. https://cursor.com/blog/secure-codebase-indexing
- 3.
Amir Al-Maamari, "Can You Trust Your Copilot? A Privacy Scorecard for AI Coding Assistants," arXiv:2509.20388, September 2025. https://arxiv.org/abs/2509.20388
- 4.
Cursor, "Ignore File," prod.cursor.com, accessed October 2026. https://prod.cursor.com/docs/reference/ignore-file
- 5.
Hacker News, "Reverse Engineering Cursor's LLM Client," news.ycombinator.com, 2025. https://news.ycombinator.com/item?id=44207063
No. Privacy Mode controls what Cursor stores and uses for training. It does not prevent code from being transmitted to AI providers like Anthropic or OpenAI during inference. Your code still reaches those servers to generate completions.
Cursor's codebase indexing stores embeddings of your files to build context for AI queries. If your repo contains hardcoded credentials in any file, those values can appear in AI context even if you never explicitly paste them. Removing credentials from files before they are indexed is the safest approach.
Yes. Terminal output such as stack traces, curl responses, and database query results often contains IPs, emails, and connection strings. Paste the terminal output into the scrubber, then paste the clean version into Cursor.
Yes. Load the page once and then disconnect from the internet. The scrubber will continue to work because all logic is contained in the JavaScript already downloaded to your browser. Nothing is fetched from a server during scrubbing.
No. CapyToolkit is an independent browser-based utility with no connection to Cursor or Anysphere. It does not interact with Cursor's API or codebase.
Remove Sensitive Data Before Sending to Grok
Grok runs on xAI's servers, which are separate from OpenAI or Anthropic infrastructure.1 Every prompt you submit to Grok via grok.com or the xAI API is transmitted to xAI's systems and processed under their data retention policy. For developers and security teams working with internal infrastructure details or customer data, this creates an exposure point that browser-based pre-scrubbing eliminates.
xAI's Grok supports real-time search and has access to X (formerly Twitter) data, making it a distinct data handling environment from other AI providers. Yet the core risk is the same: internal IPs, API keys, and personally identifiable information in your prompt leave your network the moment you click send. Replacing them with tokens before that send keeps your sensitive data where it belongs, in your browser.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →xAI's data handling environment
xAI is a separate legal entity from OpenAI, Anthropic, and Google.2 Grok submits your prompts to xAI's infrastructure for processing. Consequently, any data governance policy that restricts sending credentials to third-party AI providers covers Grok just as it covers ChatGPT or Claude. Furthermore, Grok's integration with X means its training data environment is distinct, which is a reason to be deliberate about what you submit, particularly for organizations with cross-platform data handling requirements.
xAI's terms of service and data handling policies differ from those of OpenAI, Anthropic, and Google. Treating Grok as covered by the same AI governance framework you apply to other providers is the safest assumption, even if xAI's specific policy language is less detailed. The scrubber treats all AI providers identically: it removes the same 22 data types regardless of where the prompt is headed. This provider-agnostic approach means you apply the same scrubbing discipline whether your team prefers Grok, Claude, or ChatGPT for a given task.
Detecting sensitive data before Grok queries
The scrubber catches all sensitive data types that typically appear in developer prompts: IPv4 and IPv6 addresses, API keys with specific patterns (sk-, AWS AKIA, GitHub ghp_, GitLab glpat-), JWT tokens, database connection strings, PEM private keys, and email addresses. Building on this, it detects personal identifiers like phone numbers, credit card numbers, SSNs, and IBAN bank account numbers, which are relevant for business and HR teams who use Grok for productivity tasks that touch customer or employee records.
For Grok specifically, business productivity queries carry a distinct risk profile. Teams use Grok for email drafting, document analysis, and competitive research, tasks that frequently involve customer contact details, internal project codenames, and confidential business data. The scrubber detects the structured PII patterns (emails, phone numbers, SSNs, IBANs) in these documents, including values embedded in prose paragraphs rather than in structured fields. Before pasting a customer contract or employee record into Grok, run it through the scrubber to catch identifiers that are easy to overlook in unstructured text.
Restoring Grok's response to real values
After Grok responds with tokens like [IP_1] and [EMAIL_1], paste the response into the Restore tab and upload the variables file you downloaded after scrubbing. The restoration runs locally in your browser and replaces every token with its original value in under a second. Your team receives a natural-language response that references your real infrastructure without Grok ever having seen it. Conversely, skipping the scrub means the sensitive values are already in xAI's systems before any policy controls can act.
The restore step is identical regardless of which AI provider generated the response. Whether Grok, ChatGPT, Claude, or any other tool produced the tokenized output, the same variables file reverses the tokenization. This consistency is important for teams that rotate between AI providers: the scrub-and-restore workflow does not change based on which tool you use. Download the variables file after scrubbing, and it works for any response that contains your tokens, no matter where that response came from.
xAI developer API versus grok.com: the same transmission risk
xAI provides a developer API at api.x.ai in addition to the grok.com web interface.1 Both channels carry identical data transmission risk: your prompt travels to xAI's infrastructure regardless of whether you submit through a browser or an API call. Developers building applications on the xAI API should apply the same scrubbing step to user input before constructing API requests as they would before pasting text into grok.com directly.
The API path creates a second exposure vector that developers sometimes overlook. A web team that remembers to scrub manual prompts at grok.com may not apply the same discipline to system-prompt content or user input piped programmatically through the xAI API. Your scrubbing procedure should cover both channels: scrub the system prompt content at deployment time and scrub dynamic user input at the point of API call construction.
Rate limits and token budgets do not affect data handling
Using the xAI API under a paid plan or a free tier does not change xAI's data handling obligations for your prompts. xAI states it does not use business inputs or outputs for training its models, yet those inputs and outputs are automatically deleted only within 30 days unless retention is agreed in writing or the content is legally required to be kept.3 Review xAI's data processing documentation and terms of service for your specific tier before making assumptions about retention; when in doubt, scrub all sensitive values in every API request regardless of tier.
Grok's real-time X search and prompt content influence
Grok integrates with X (formerly Twitter) to retrieve real-time information in its responses.4 When you submit a prompt, Grok may generate search queries from it and retrieve matching X content to supplement its answer. The generated queries derive from the content of your prompt, not just your explicit search terms. A prompt containing an organization's name, a product name, or a specific customer's company name can influence what Grok searches for on X.
Scrubbing organizational identifiers from your prompt reduces the chance that Grok constructs X search queries referencing your internal specifics. The structural question in your prompt remains intact after scrubbing; only the identifying context is removed. Grok can analyze a configuration structure, debug a connection pattern, or explain an error format without knowing which specific organization or product is involved.
Grok and public X data as context for your prompt
Grok can surface public X posts that mention the names, companies, or technical terms in your prompt. For organizations with a public social media presence, a prompt containing your company name may cause Grok to retrieve and reference public X posts about your company as context for its response. Strip organizational identifiers before Grok searches X, since removing company names and other identifiers from prompts before submission prevents your query from triggering content retrieval that surfaces information you did not intend to bring into the analysis.
Business productivity tasks with Grok and organizational data
Business teams use Grok for productivity tasks beyond technical debugging: email drafting, meeting summarization, contract review, and competitive research. Each of these tasks can involve customer contact information, employee records, or confidential negotiation details that fall under your organization's data handling policy. The volume of sensitive data flowing through these productivity workflows often exceeds what teams handle in developer debugging sessions, because business users process full documents rather than targeted code snippets, and a single contract review can expose dozens of identifiable values in one paste.
Because these workflows move real personal data into a third-party provider, they sit in a specific regulatory context, including the European Data Protection Board's position that personal data disclosed to AI providers is processed by that provider.5 This means GDPR, CCPA, or internal governance policies apply to what you send to Grok.
The two-step scrub-and-restore workflow applies to business productivity tasks the same way it applies to developer debugging. Scrub the customer record or contract text before submitting, receive Grok's structural analysis or draft response with tokens in place, then restore the real values using the variables file. Your team reads the final output with real names and contact details without those values ever having reached xAI's servers. The workflow adds under a minute to each task while eliminating the compliance exposure entirely, which is a worthwhile trade for documents containing regulated personal data subject to GDPR, CCPA, or internal governance policies.
Scrubbing before Grok API integrations in business applications
Business applications that integrate the xAI API for document analysis, customer communication assistance, or data summarization inherit the same transmission risk as direct grok.com usage. Scrub user-submitted content at the API boundary before it is included in any request to the xAI API. For applications that process customer records, add a scrubbing step to the data preparation pipeline rather than relying on individual contributors to scrub manually at the point of prompt construction. This pipeline-level approach ensures consistent protection regardless of which team member initiates the request.
Applying the same scrub step to both the xAI API boundary and the grok.com interface keeps protection consistent across every entry point. CapyToolkit runs the scrubber locally in your browser, so no prompt text is uploaded to any server, and the variables file is the only record tying the tokens back to the real customer data.
When to use this
Use this before submitting any prompt to Grok that contains internal network details, credentials, customer records, or any data your organization's AI usage policy restricts from third-party transmission.
Examples
Security analysis with internal IP ranges
Analyze this firewall rule: ALLOW 192.168.10.0/24 → 10.0.0.1 port 443. Our WAF is at waf.corp.internal.
Analyze this firewall rule: ALLOW [IP_1]/24 → [IP_2] port 443. Our WAF is at [DOMAIN_1].
Customer data summary for a productivity task
Summarize the situation: customer [email protected] called about invoice #12345, card ending in 4242.
Summarize the situation: customer [EMAIL_1] called about invoice #12345, card ending in [CC_1].
Review the variables file to confirm the card pattern match is correct.
- 1.
xAI, "X Search," docs.x.ai, accessed June 2026. https://docs.x.ai/developers/tools/x-search
- 2.
xAI, "Tools Overview," docs.x.ai, accessed June 2026. https://docs.x.ai/developers/tools/overview
- 3.
xAI, "Enterprise FAQs," x.ai, February 2025. https://x.ai/legal/faq-enterprise
- 4.
Kyle Wiggers, "xAI's Grok chatbot can now 'see' the world around it," TechCrunch, April 2025. https://techcrunch.com/2025/04/22/xais-grok-chatbot-can-now-see-the-world-around-it/
- 5.
European Data Protection Board, "Report of the work undertaken by the ChatGPT Taskforce," edpb.europa.eu, May 2024. https://www.edpb.europa.eu/documents/task-force-report/report-of-the-work-undertaken-by-the-chatgpt-taskforce_en
Grok is developed by xAI, a company closely associated with X. Grok has access to X's data for real-time information. For data governance purposes, treat Grok as a distinct third-party AI provider. Your prompts reach xAI's servers, separate from your own infrastructure.
xAI's data handling and model training practices are governed by their terms of service, which can change. Regardless of current policy, scrubbing sensitive values before submission ensures that no identifiable data reaches xAI's systems.
Around 22 types, including IPv4 and IPv6 addresses, API keys (multiple vendor formats), JWT tokens, database connection strings, PEM private keys, email addresses, credit card numbers, SSNs, phone numbers, and IBAN bank account numbers.
The scrubber works on any text, including API prompts, browser messages, or file content. Copy the text you plan to use as a Grok prompt, paste it into the scrubber, copy the cleaned version, then use that cleaned version in your API call or browser session.
No. CapyToolkit is an independent browser utility with no connection to xAI, Grok, or X. It does not access or interact with any AI provider's API.
Remove Sensitive Data Before Querying Perplexity
Perplexity combines large language model reasoning with real-time web search. When you submit a query, Perplexity sends it to its servers, runs web searches based on your question, and synthesizes results using an LLM, all in a single round-trip that includes the full text of your prompt. A query containing internal credentials or customer data travels to Perplexity's infrastructure, its search providers, and the underlying model.1
Unlike pure chat AI tools, Perplexity's architecture means your prompt can influence what URLs are fetched and what external sources are queried. Removing sensitive values before the prompt is submitted ensures that neither Perplexity's servers nor any third-party search infrastructure receive identifiable data from your organization.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →How Perplexity processes your query
Perplexity's pipeline takes your prompt, generates search queries from it, retrieves web content, and passes both the original prompt and the retrieved content to an LLM for synthesis.2 Consequently, a prompt that includes a customer's email address or an internal API key could influence what search terms Perplexity sends to the open web. Scrubbing before submission closes this channel entirely: the search generation step receives only tokens, which produce general, non-identifying queries.
When you paste a technical error message into Perplexity for debugging help, the pipeline may generate search queries based on the specific hostnames, IP addresses, and credential patterns in that error. Those search queries go to Perplexity's search providers, creating a trail that references your internal infrastructure. After scrubbing, the error message retains the structural information Perplexity needs for analysis (error type, stack trace shape, protocol details) while the search queries it generates contain only generic terms like [IP_1] and [DOMAIN_1] that reveal nothing about your environment.
What to scrub before a Perplexity query
For research and productivity queries, the most common sensitive elements are email addresses, phone numbers, company names embedded in contract text, and personal identifiers in HR or legal documents. For technical queries, internal IP addresses, API keys, and database hostnames often slip into error messages or stack traces that you paste for context. Furthermore, IBAN and credit card numbers can appear in finance team queries. The scrubber detects all 22 types so you can paste freely without auditing each field manually.
You should treat every paste into Perplexity as a potential data transmission event, because the search-augmented pipeline means your prompt content can influence external web queries.3 Before pasting a legal document, scrub the party names, email addresses, and account numbers. Before pasting a technical log, scrub the internal IPs and hostnames. The scrubber handles all of these in a single pass, so the additional step adds minimal time to your workflow while eliminating the risk of your prompt content reaching search providers in identifiable form.
Pro plan and data handling
If you use Perplexity Pro, you can disable conversation history, which reduces local logging. Yet the prompt itself is still processed server-side on every request, regardless of history settings. There is no true local processing mode. Building on this: even for one-off queries with no stored history, the transmission window exists. Pre-scrubbing before any Perplexity query, whether on Pro, free, or API, eliminates the window regardless of account settings.
You should not rely on account-level settings as your primary data protection control. History settings control what Perplexity stores after the fact, but they do nothing to prevent the initial transmission of your prompt to Perplexity's servers and search providers. Scrubbing before submission is the only control that operates before the transmission event. Apply it consistently across all Perplexity tiers, because the transmission risk is identical whether you use the free tier, Pro, or the API.4
Perplexity Enterprise and organizational data controls
Perplexity Enterprise is the tier designed for organizations with data governance requirements. It offers zero data retention on inputs and outputs, no model training on organizational prompts, and admin controls for managing user access and feature availability. This is distinct from Perplexity Pro, which only disables conversation history logging; the Pro tier does not provide zero data retention or training exclusion.5
For teams using the standard free tier or Pro, no organizational data controls are available beyond the individual history setting. Scrubbing before submission is the only available technical control that prevents sensitive values from reaching Perplexity's infrastructure. For teams evaluating Perplexity Enterprise, scrubbing remains the correct approach during the evaluation period before a formal Enterprise agreement is in place.
Verifying your Perplexity account tier
Your subscription tier appears under Account Settings in the Perplexity web interface. The history disable toggle appears under the same settings. If your organization has a Perplexity Enterprise agreement, confirm with your admin that zero retention is enabled at the team level before reducing your scrubbing discipline. For all tiers, scrubbing is the technical control that works independently of account configuration.
How Perplexity constructs search queries from your prompt
Perplexity's search-augmented pipeline takes your prompt, derives search queries from it, retrieves web content, and synthesizes a response. The search queries generated from your prompt are influenced by the entities and concepts in your text. A prompt containing your company's internal project name, a client's company name, or a specific technical product may cause Perplexity to generate search queries that reference those names in public search index requests.
Scrubbing identifying names and context from your prompt prevents this query inference from surfacing public web information associated with your specific organization. The goal is to keep the structural question, drop the org name: a legal framework, a technical approach, or a policy interpretation survives scrubbing intact, and Perplexity retrieves relevant sources for the general question without knowing which specific organization or individual it relates to.
Source citation and traceability for scrubbed queries
Perplexity cites sources in its responses. When you scrub identifying context from a prompt, the sources Perplexity cites reflect general search results for the structural question rather than results specific to your organization. For research tasks where source quality matters, review the citations in Perplexity's response before restoring tokens: the cited sources should address the general domain of your question, and you then apply the restored context to those findings independently.
Using the Perplexity API in developer applications
The Perplexity API follows the same transmission pipeline as the web interface: prompts travel to Perplexity's infrastructure, search queries are generated, web sources are retrieved, and the response is synthesized server-side. Developers building applications on the Perplexity API should apply scrubbing to user input and system prompt content before constructing API requests, exactly as they would for the web interface.
Applications that use Perplexity for document analysis, customer research, or knowledge base queries often include user-submitted text as part of the prompt. Scrub user input at the API boundary rather than relying on users to scrub before submitting to your application. This approach provides a consistent and automatic scrubbing layer regardless of whether individual users are aware of the data handling implications of their submissions.
Restoring Perplexity responses when tokens appear in citations
When Perplexity includes tokens from your scrubbed prompt in its response (referencing [IP_1] or [EMAIL_1] directly in its answer), paste the full response into the Restore tab alongside your variables file. The restoration replaces all tokens with original values simultaneously. Citations that Perplexity includes from web sources are not affected by the restore step, since they contain general text rather than your specific tokens.
Restoring after the response means the web sources Perplexity cites stay general, while your real identifiers return only in your own browser. Because CapyToolkit performs the replacement locally, no restored value is transmitted back to Perplexity, and the scrubbed prompt you originally sent already contained no personal data for the search pipeline to process.
When to use this
Use this before submitting any Perplexity query that contains customer data, internal network details, personal identifiers from documents, or credentials that your organization restricts from third-party services.
Examples
Research query with customer context
What are my options if customer Jane Smith ([email protected], SSN 123-45-6789) wants a refund under California law?
What are my options if customer [EMAIL_1] ([SSN_1]) wants a refund under California law?
Perplexity can research the legal question without needing the real customer identity.
Technical query with internal infrastructure details
Why might a request from 10.0.1.44 to api.prod.internal port 8443 fail intermittently?
Why might a request from [IP_1] to [DOMAIN_1] port 8443 fail intermittently?
- 1.
The AI Engineer, "How Perplexity Built Their Search Engine," substack.com, June 2026. https://theaiengineer.substack.com/p/how-perplexity-built-their-search
- 2.
Perplexity AI, "Privacy Policy," perplexity.ai, February 2026. https://www.perplexity.ai/hub/legal/privacy-policy
- 3.
OWASP Foundation, "LLM02:2025 Sensitive Information Disclosure," owasp.org, 2025. https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/main/2_0_vulns/LLM02_SensitiveInformationDisclosure.md
- 4.
European Data Protection Board, "Report of the work undertaken by the ChatGPT Taskforce," edpb.europa.eu, May 2024. https://www.edpb.europa.eu/documents/task-force-report/report-of-the-work-undertaken-by-the-chatgpt-taskforce_en
- 5.
Perplexity Security Team, "How Perplexity Enterprise Pro Keeps Your Data Secure," perplexity.ai, April 2025. https://www.perplexity.ai/hub/blog/how-perplexity-enterprise-pro-keeps-your-data-secure
Perplexity generates search queries based on your prompt and retrieves web content. The raw text of your prompt is processed on Perplexity's servers, not sent verbatim to search engines. Regardless, the prompt text reaches Perplexity's infrastructure, which is a transmission event your data governance policy may restrict.
It reduces logging of conversations on Perplexity's side. It does not prevent your prompt from being transmitted and processed server-side during the request. Scrubbing before submission provides protection that history settings cannot.
Around 22 types: IPv4 and IPv6 addresses, email addresses, API keys (multiple formats), AWS keys, GitHub and GitLab tokens, JWT tokens, database connection strings, PEM private keys, credit card numbers, SSNs, phone numbers, and IBAN bank account numbers.
Yes. The scrubber works on any text. Copy your planned prompt, paste it into the scrubber, copy the clean version, then include it in your API call body. The two-step workflow applies regardless of how you access Perplexity.
No. CapyToolkit runs entirely in your browser. No text is transmitted. You can confirm this by opening DevTools and watching the Network tab while you paste and scrub.
Remove Sensitive Data Before Using Notion AI
Notion AI processes workspace content on Notion's servers. When you invoke an AI feature such as summarization, Q&A, or writing assistance, Notion sends the selected pages, blocks, or your explicit prompt to its AI provider for processing. Any PII, credentials, or internal data embedded in those pages travels off your Notion workspace to external AI infrastructure.1
Teams that use Notion to store HR data, product roadmaps with internal hostnames, or customer-facing records need a scrubbing step before using Notion AI features on those pages. Removing sensitive fields before pasting them into AI-assisted blocks or prompts ensures that Notion's AI pipeline sees only tokenized placeholders.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →How Notion AI accesses your workspace content
Notion AI uses context from the current page and your explicit prompt to generate responses. For team wikis and databases, this means that invoking Notion AI on a page containing employee records, customer contracts, or infrastructure runbooks sends that page's content to Notion's AI provider. Consequently, a team that documents internal IP addresses and SSH credentials in a Notion runbook risks exposing those values every time an AI feature is triggered on that page. Pre-scrubbing the source content before it enters Notion removes this risk.
This passive context sharing is the primary difference between Notion AI and paste-based AI tools. With ChatGPT or Claude, you choose what to paste. Notion AI reads surrounding page content automatically, which means sensitive data you did not intend to share can enter the AI context without an explicit copy-paste action.2 Before you type any AI prompt on a page that contains PII, check the surrounding content for identifiers. If the page contains sensitive data, scrub it at the source before invoking AI features.
Sensitive data patterns in Notion content
Notion pages typically contain a mix of text, inline code, and database fields. The scrubber handles all text content and detects personal identifiers (email addresses, phone numbers, SSNs, IBANs), infrastructure data (IPv4 and IPv6 addresses, internal domain names), credentials (API keys in multiple formats, database connection strings, JWT tokens), and payment data (credit card numbers). Building on this, it processes multi-line input, making it suitable for pasting entire database exports or documentation pages for rapid scrubbing before content is entered into Notion.
You can use the scrubber to audit your existing Notion content before enabling AI features. Copy the text of your most sensitive pages (employee databases, client records, infrastructure documentation) and paste them into the scrubber. The variables file shows you exactly which identifiers exist in each page, giving you a PII inventory for your Notion workspace. Use this inventory to decide which pages need scrubbing before you enable AI features on them, and which pages are safe to use with AI without modification.
Enterprise Notion and AI data handling
Notion Enterprise offers AI controls that allow admins to disable AI features organization-wide.3 Yet teams with AI enabled, which is the default for most paid plans, need a workflow to use those features safely on sensitive content. Furthermore, the AI provider receives the content for processing; Notion's own data retention policy does not govern what the AI provider does with submitted data under their enterprise agreements. A pre-scrub workflow makes Notion AI safe for HR teams, legal, and IT without requiring org-level feature disablement.
Enterprise admins face a trade-off: disabling AI features organization-wide protects sensitive data but removes a productivity tool that teams rely on for non-sensitive work. A more practical approach is to enable AI features but restrict them to pages that have been scrubbed or verified as PII-free. Create a team convention where pages containing sensitive data are tagged or placed in a specific database, and train your team to check for that tag before invoking AI. This lets teams benefit from Notion AI on general knowledge-base pages while maintaining a hard boundary around sensitive content.
Notion AI context scope: pages, linked databases, and block context
Notion AI's context includes more than just your typed question. When you invoke AI on a page, Notion sends the surrounding page content, any linked database views visible on that page, and the blocks you have selected as context for the AI request. Invoking "Ask AI" in a page that embeds a filtered customer database view sends the visible database records as part of the request context, even if you typed a question unrelated to those records.4
For pages that embed database views containing PII columns, either filter the database view to exclude sensitive columns before invoking AI, or scrub the data before it enters the Notion page. Creating a separate scrubbed-content page for AI analysis, copied from the original source, gives you a stable workspace for AI tasks without exposing the source records directly. This approach separates the original PII-containing database from the AI-assisted workflow layer.
Inline AI commands and surrounding block context
Inline AI commands (typing "/" and selecting an AI option within a block) read the surrounding page content for context. A simple "Improve writing" command on a paragraph in a page that also contains an HR table with SSNs and IBANs sends that surrounding context to Notion's AI provider. Invoke inline AI commands only on pages that contain exclusively non-sensitive content, or scrub the page before any inline AI command if the surrounding block still holds an HR table or other sensitive data.
Meeting notes and recurring PII exposure in team wikis
Team meeting notes in Notion frequently follow templates that include attendee email addresses, action-item owner names, and client names as structured fields. These notes accumulate in shared databases and are queried via Notion AI for project summaries, sprint retrospectives, and historical reference. Each AI query on a meeting notes database sends the records in the query scope to Notion's AI provider as context.
Replace individual email addresses with role names in meeting note templates to reduce PII density at the source. Instead of logging "[email protected] assigned to fix the auth bug," use "Auth lead assigned to fix the auth bug." This structural change reduces the amount of PII in meeting notes without losing the operational information teams need, and it makes Notion AI features safe to use on the notes database without a scrubbing step at query time.
Client names and external company references in Notion projects
Project pages that reference client company names, client contacts, and contract details create a PII exposure when Notion AI is used for project analysis. Client company names combined with project status and financial context constitute business-sensitive information that your organization may restrict from third-party AI providers under NDA or client agreement terms. Scrub client identifiers before invoking AI on project pages to satisfy these obligations regardless of Notion's AI provider arrangements.
Notion guest access and AI processing scope
Notion workspaces with guest access allow external collaborators to view and edit pages in your workspace. When AI features are enabled at the workspace level, guests with editor access can invoke Notion AI on pages they can access. If those pages contain internal employee data or client records, guests can trigger AI processing of that sensitive content through Notion's AI pipeline without additional authorization.
Review guest permissions for workspaces where Notion AI is enabled and where sensitive data exists in the same page space as guest-accessible content. Restrict sensitive databases to full members rather than guests, or create guest-specific pages that contain only the non-sensitive context that external collaborators need. Separating the permission scope of sensitive content from the permission scope of guest-accessible content prevents guests from inadvertently triggering AI processing on data they should not be able to submit externally.
Notion AI provider chain and DPA coverage
Notion AI is powered by third-party AI providers, including OpenAI.5 When you invoke Notion AI, the content Notion sends to its AI providers is covered by Notion's data processing agreements with those providers, not by your own DPA with the AI provider directly. Review Notion's AI data handling documentation for your plan tier to understand which providers receive your content and under which retention terms. Scrubbing before submission ensures that the content transmitted through this provider chain contains no identifiable values, regardless of the specifics of Notion's backend arrangements.
Scrubbing before you invoke Notion AI means the content reaching OpenAI through that provider chain already carries no identifiable values. CapyToolkit runs the scrubber in your browser, so the sensitive text never leaves your device during processing, and the variables file remains under your own control rather than inside Notion's infrastructure.
When to use this
Use this before invoking any Notion AI feature on a page, database view, or document that contains HR records, internal infrastructure details, customer identifiers, or credentials.
Examples
Runbook page with internal endpoints
Restart procedure: SSH to 10.20.1.15, run sudo systemctl restart app. DB at postgres://ops:[email protected]/prod.
Restart procedure: SSH to [IP_1], run sudo systemctl restart app. DB at [DBURL_1].
Paste the scrubbed version into Notion before using AI summarization or Q&A on it.
HR database export with employee records
Name: Alice Brown, Email: [email protected], Phone: (415) 555-0192, Salary: $112,000
Name: Alice Brown, Email: [EMAIL_1], Phone: [PHONE_1], Salary: $112,000
- 1.
Notion, "Notion AI security & privacy practices," notion.com, accessed June 2026. https://www.notion.com/help/notion-ai-security-practices
- 2.
OWASP Foundation, "LLM02:2025 Sensitive Information Disclosure," owasp.org, 2025. https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/main/2_0_vulns/LLM02_SensitiveInformationDisclosure.md
- 3.
NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0)," nist.gov, January 2023. https://www.nist.gov/itl/ai-risk-management-framework
- 4.
European Data Protection Board, "Report of the work undertaken by the ChatGPT Taskforce," edpb.europa.eu, May 2024. https://www.edpb.europa.eu/documents/task-force-report/report-of-the-work-undertaken-by-the-chatgpt-taskforce_en
- 5.
Notion, "Notion's Commitment to AI Safety," notion.com, accessed June 2026. https://www.notion.com/help/ai-safety
Yes. Notion AI uses OpenAI's API for generation. When you invoke an AI feature, Notion sends the relevant content to OpenAI's servers for processing. Notion's data processing agreements with OpenAI govern how that content is handled, not your own privacy policy.
Notion Enterprise allows workspace admins to disable AI features entirely. For teams on lower tiers or where AI features are needed for non-sensitive content, a pre-scrub workflow is a practical middle ground.
The scrubber handles any text format. For database exports, paste the text representation (CSV, JSON, or plain copied table content) into the scrubber, clean it, then paste the result into Notion. Field names and non-sensitive values are preserved.
Nothing outside your browser. CapyToolkit does not transmit, store, or log any text you paste. All processing runs locally using JavaScript in your browser tab.
Both. Database records can contain email addresses, phone numbers, and other PII. Pages can contain infrastructure details, credentials in code blocks, and employee data. The scrubber works on any text content copied from Notion.
Remove Secrets from Code Before Pasting to AI Assistants
Developers paste code into AI assistants dozens of times a day for debugging, code review, and explanations. When that code contains hardcoded credentials, internal hostnames, or private keys left in the repo, those secrets travel to the AI provider's servers.
This happens routinely: a developer pastes a failing config file to ask why the connection is refused, and the database password goes to OpenAI with it. The scrubber detects all major credential formats before the paste reaches the AI, replacing each with a token that preserves the structural question without disclosing the secret.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →Where credentials hide in code
Credentials appear in code in predictable places: .env files, docker-compose.yml environment sections, deployment scripts, CI/CD pipeline configs, and test fixtures. Less obviously, they appear in stack traces (connection strings in error messages), git commit context (API keys accidentally committed then removed), and in-code comments left by developers documenting configuration. Consequently, a developer pasting "just a config file" for debugging help routinely discloses production database passwords, payment API keys, and internal service credentials to the AI provider.
The risk is highest in organizations without a pre-commit scanning tool installed. A developer who accidently commits a Stripe key and then force-pushes to remove it from history still faces exposure if that key was briefly visible on a shared branch. When that developer pastes the failing CI log into an AI tool to debug the pipeline failure, the scrubber catches the credential patterns in the log output. Without the scrubber, the log text containing the live key travels to the AI provider as part of the pasted context, creating a second exposure event from the same original mistake.
Credential formats detected in code
The scrubber detects all major developer credential types: generic API keys (sk- prefix), AWS access keys (AKIA), GitHub PATs (ghp_), GitLab tokens (glpat-), Stripe keys (sk_live_, sk_test_, pk_live_, pk_test_, rk_live_), Google API keys (AIza), SendGrid keys (SG.), Twilio Account SIDs (AC + 32 hex), NPM auth tokens (npm_), Slack tokens (xox*), JWT tokens (three-segment eyJ header), database connection strings (postgres://, mysql://, mongodb://, redis://), and PEM private key blocks. Building on this, it catches IPv4 and IPv6 addresses, internal hostnames, and email addresses, all common in stack traces and log output pasted for debugging.
Vendor-specific patterns make detection highly precise. An AKIA-prefixed string is almost certainly an AWS access key, and a ghp_-prefixed string is almost certainly a GitHub token. The generic sk- prefix casts a wider net and may occasionally flag non-credential strings that happen to match. After scrubbing, review the variables file to confirm that every flagged value is genuinely a credential. This review takes seconds and gives you confidence that the scrubbed output contains no secrets before it reaches any AI coding assistant.
The AI can still help with scrubbed code
An AI coding assistant that sees [API_1] instead of sk-prod-key-xyz123 understands that [API_1] is a placeholder for a secret value. It can still identify security issues in the code, suggest a better configuration structure, explain the connection logic, and review the surrounding code, but it just never learns your real credential. Conversely, providing the real key gives the AI no additional ability to help with the code question, while creating a transmission event that your data governance policy likely prohibits.
The structural information in code is what makes AI assistance valuable, not the specific values. A connection string pattern (postgres://user:password@host:port/db) teaches the AI everything it needs to know about your database architecture. The actual password in that string adds zero diagnostic value. Scrubbing preserves the architectural signal while removing the secret, giving you the same quality of AI assistance with none of the credential exposure. For most code review and debugging tasks, the scrubbed version and the original version produce identical AI output.
Infrastructure-as-code files and CI/CD credential exposure
Infrastructure-as-code files carry the highest credential density of any developer file type. Terraform .tfvars files store AWS access keys, database passwords, and API endpoints as plain text variables. Kubernetes Secret manifests in YAML encode credentials in base64 (which is not encryption) and appear in code review contexts. GitHub Actions workflow files, GitLab CI .gitlab-ci.yml, and Jenkins Jenkinsfile configurations accumulate env: blocks with production secrets that developers paste into AI assistants when debugging pipeline failures.
Scrubbing before pasting any IaC file removes credentials in a single pass without requiring you to identify each sensitive field manually. The scrubber detects all major API key formats, database connection string schemas, and IPv4 addresses that commonly appear in Terraform variable files and CI/CD configs. After scrubbing, the AI assistant can analyze your pipeline logic, suggest structural improvements, and review your workflow configuration without ever receiving a credential.
Handling Terraform state files
Terraform state files (terraform.tfstate) contain the actual resource values applied by Terraform, including resource IDs, connection strings, and sometimes credentials that were passed as inputs. State files are often pasted into AI tools for debugging resource drift or import issues; scrub Terraform state before debugging drift, since the AI can analyze resource relationships and state structure from the tokenized version just as effectively as from the original.
AI coding assistant context windows and passive credential capture
AI coding assistants read more than what you explicitly paste. GitHub Copilot reads the active file, recently edited files, and adjacent files in your project.1 Cursor's Composer and Codebase Indexing features read your entire repository to build context.2 When a .env file or secrets.yml file is open in your editor, its content is eligible to enter the context window that the AI coding assistant sends upstream, even if you never explicitly paste it.
Close or unload sensitive files from your editor before starting an AI-assisted coding session on unrelated code. Alternatively, use a .gitignore-style exclusion pattern in Copilot Business or Cursor's Privacy settings to prevent specific file patterns from entering the context. The scrubber addresses the paste-based exposure; editor-level content exclusion addresses the passive context capture that occurs without an explicit paste action.
Pre-scrub as a habit before every paste to an AI coding tool
The most reliable protection is scrubbing everything before it enters any AI coding context, regardless of whether it appears sensitive at first glance. Config files that look harmless often contain staging credentials or internal hostnames embedded in template values. Establishing a workflow of paste-into-scrubber-first eliminates the need to make case-by-case judgments about whether a file contains sensitive content. The scrub takes seconds and the habit prevents the entire category of accidental credential disclosure.
Steps to take after a credential reaches an AI tool
Revoke the exposed credential immediately. Assume the AI provider's infrastructure has seen it, regardless of the provider's stated privacy policy, because the transmission event already occurred. Generate a new credential, update all systems that used the old one, and confirm the new credential is working before decommissioning the old one. Do not wait to see if an incident occurs before rotating; credential rotation after exposure is mandatory, not optional.
Document the incident for your security team even if no visible harm resulted. AI provider data retention means the credential may persist in server logs beyond the session. For credentials with broad permissions (AWS root keys, GitHub organization tokens, Stripe live keys), notify your security team and treat the incident as a credential breach requiring a formal record. Many compliance frameworks (SOC 2, ISO 27001) require logging credential exposure events regardless of whether exploitation was observed.34
Preventing future exposure with pre-commit hooks
Pre-commit hooks using tools like git-secrets, truffleHog, or detect-secrets scan staged files for credential patterns before each commit.5 These tools complement the scrubber: the scrubber prevents credentials from leaving your machine in paste-based AI workflows, while pre-commit hooks prevent credentials from entering your repository in the first place. Running both controls closes the two most common paths for credential exposure in developer workflows.
The scrubber complements these checks by catching secrets at the moment they enter an AI prompt, which is the exposure path pre-commit hooks cannot reach. Because CapyToolkit processes everything in your browser with no upload step, the two controls together cover the credential lifecycle from repository to runtime without adding a separate server component to maintain.
When to use this
Use this before pasting any code snippet, config file, log output, or stack trace into GitHub Copilot, Claude, ChatGPT, Cursor, or any other AI coding assistant.
Examples
Config file with credentials
DATABASE_URL=postgres://admin:[email protected]:5432/mydb AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE STRIPE_KEY=sk_live_abcdefghijk1234567890
DATABASE_URL=[DBURL_1] AWS_ACCESS_KEY_ID=[AWS_1] STRIPE_KEY=[STRIPE_1]
Stack trace with internal IPs
ConnectionError: Failed to connect to 10.20.30.40:6379 Host: redis.internal.company.com
ConnectionError: Failed to connect to [IP_1]:6379 Host: [DOMAIN_1]
- 1.
GitHub, "Best practices for using GitHub Copilot to work on tasks," docs.github.com, accessed June 2026. https://docs.github.com/en/copilot/tutorials/cloud-agent/get-the-best-results
- 2.
Cursor, "Securely indexing large codebases," cursor.com, 2025. https://cursor.com/blog/secure-codebase-indexing
- 3.
Wikipedia, "ISO/IEC 27001," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/ISO/IEC_27001
- 4.
Wikipedia, "System and organization controls," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/System_and_Organization_Controls
- 5.
OWASP Foundation, "Secrets Management Cheat Sheet," cheatsheetseries.owasp.org, accessed October 2026. https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
All major credential formats: generic API keys (sk- prefix), AWS access keys (AKIA), GitHub tokens (ghp_), GitLab PATs (glpat-), Stripe keys (sk_live_, sk_test_, pk_live_), Google API keys (AIza), SendGrid keys (SG.), Twilio Account SIDs (AC + 32 hex chars), NPM auth tokens (npm_), Slack tokens (xox*), JWT tokens, and full database connection strings (postgres://, mysql://, mongodb://, redis://). PEM private key blocks, IPv4 and IPv6 addresses, internal hostnames, emails, credit card numbers, US Social Security Numbers, phone numbers, and IBAN bank account numbers are also detected.
At minimum, use it whenever you paste a file that could plausibly contain secrets such as config files, .env examples, deployment scripts, or log output with connection strings. It takes seconds to scrub and eliminates a risk category entirely.
Yes. The AI understands that [API_1] is a placeholder for a secret value. It can still identify security issues, suggest fixes, and explain the code without ever learning your real credentials.
No. CapyToolkit is a proprietary browser-based utility, and the scrubber tool runs entirely client-side to remove secrets from your code before pasting into AI tools. You can verify the local-only processing by opening DevTools and watching the Network tab while you scrub.
Both. For browser-based AI tools, paste the scrubbed text directly. For API-based tools (Copilot, Cursor plugins), copy the scrubbed text and use it as the prompt content in your API call.