PII Scrubbing for GDPR, CCPA, HIPAA and PCI DSS
Privacy law rarely bans AI tools outright. What it regulates is the moment personal data leaves your control: GDPR asks whether you had a lawful basis and a processor contract, CCPA whether the disclosure counts as a sale or share, HIPAA whether a Business Associate Agreement covers the recipient, and PCI DSS whether the system that receives card data sits inside your assessed scope. Remove the regulated values before the prompt is sent, and most of those questions never come up for that prompt.
A local scrubber handles the part of that job that follows a pattern. It replaces emails, phone numbers, card numbers, SSNs, IBANs, IP addresses and credentials with tokens in your browser, and it leaves names, street addresses and dates for you to remove. The sections below cover what each framework treats as regulated, what the scrubber catches for it, and where a contract or a certified process still has to do the rest.
Before regulated text reaches an AI tool
- Pattern identifiers emails, phone numbers, card numbers, SSNs, IBANs and IP addresses are tokenized in one pass
- Names, addresses and dates remove them by hand, since no fixed pattern identifies them
- Short card codes CVVs and PINs are too short to detect reliably, so delete them yourself
- Variables file store it as carefully as the original data, because it maps every token back to a real value
Opens the PII Scrubber with this page's checklist shown at the top of the tool.
Open in the tool →What the four frameworks have in common
Each framework uses its own vocabulary, yet all four reward the same habit: send the smallest amount of regulated data that still gets the job done. Scrubbing is a direct way to practise that habit at the prompt level, and the overlap between the frameworks explains why one tool can serve all four without a separate configuration for each.
Minimization is the shared principle
GDPR states it most plainly. Personal data must be adequate, relevant and limited to what is necessary for the purpose, which Article 5 calls data minimisation.1 HIPAA's Minimum Necessary standard, CPRA's right to limit the use of sensitive personal information and PCI DSS's rule against keeping card data you do not need all point in the same direction, toward less regulated data in fewer places.
A drafting or debugging prompt rarely needs the real email address, card number or patient phone number to produce a useful answer. Replacing those values with tokens keeps the structure the AI needs and drops the identifier the law cares about, which is minimization applied at the exact point where data would otherwise leave your device.
Where scrubbing stops and process takes over
Pattern matching cannot certify anything. HIPAA's Safe Harbor method lists 18 identifiers that must all be removed before health data counts as de-identified, including names, geographic subdivisions smaller than a state and most dates tied to a person.2 The scrubber detects several of those by pattern and none of the free-text ones, so a scrubbed clinical note is safer but not de-identified in the legal sense.
The same limit applies everywhere. Treat the scrubber as the first pass that removes the identifiers you would otherwise miss under time pressure, then read the output for names and addresses, and keep the contracts your framework requires with any AI provider that may still receive regulated data, such as a Business Associate Agreement under HIPAA or a processor agreement under GDPR.
- 1.
GDPR, "Art. 5 – Principles relating to processing of personal data," gdpr-info.eu, accessed October 2026. https://gdpr-info.eu/art-5-gdpr/
- 2.
HHS OCR, "Guidance Regarding Methods for De-identification of Protected Health Information," hhs.gov, accessed October 2026. https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html
GDPR-Ready PII Scrubber: Remove Personal Data Before AI
GDPR fines hit €7.1 billion cumulatively by 20251, and enforcement in 2026 has expanded to AI system interactions. Sending personal data about EU residents to third-party AI providers without a data processing agreement, without a lawful basis, or across an unauthorized border is a GDPR violation, one that organizations trigger routinely by pasting prompts into ChatGPT, Copilot, or Claude.
Removing EU personal data before it reaches an AI provider eliminates the most common data transmission violation entirely. The scrubber processes text locally, so no cross-border transfer occurs at the scrubbing step, and it replaces personal identifiers with tokens that carry no identifying information under GDPR's definition of personal data.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →GDPR-relevant data the scrubber detects
Under GDPR, personal data includes any information that can directly or indirectly identify an individual, and this broad definition covers far more data types than most teams initially realize. The scrubber targets the most commonly exposed types in AI prompt workflows: email addresses and phone numbers as direct identifiers, IBANs and credit card numbers as financial identifiers, SSNs as national identifiers, IP addresses as indirect identifiers confirmed as personal data under GDPR by the CJEU in Breyer v. Germany2, internal domain names as organizational identifiers that can be linked to specific individuals, and JWT tokens as session identifiers tied to authenticated users. Replacing each of these with a token before AI submission prevents any identifiable data from leaving the processing jurisdiction, which eliminates the cross-border transfer concern entirely.
Direct versus indirect identifiers under GDPR
Direct identifiers, such as email addresses and phone numbers, uniquely identify a person without additional context. Indirect identifiers, such as IP addresses or internal hostnames, require supplementary information to link back to an individual but still qualify as personal data under GDPR. The scrubber treats both categories with equal priority, tokenizing indirect identifiers alongside direct ones to ensure comprehensive de-identification before any prompt reaches an external AI provider.
GDPR Articles relevant to AI prompt submission
Article 5 requires data minimization3, meaning personal data should not be processed beyond what is strictly necessary for the stated purpose. Sending full customer records to an AI tool just to get a date formatted is a clear example of excess processing that violates this principle. Article 46 governs transfers to third countries, so if your AI provider servers are in the US, transmitting EU personal data requires a valid transfer mechanism such as Standard Contractual Clauses. Building on this, Article 28 requires a data processing agreement with any vendor that processes personal data on your behalf4, and most public AI tools do not qualify as GDPR-compliant processors under this article. Scrubbing before transmission sidesteps all three article obligations at the prompt level because the AI provider receives only non-identifying tokens.
Local processing and the data processing agreement question
When you use the scrubber, a local browser tool that never transmits your text to any server, no data processing agreement with CapyToolkit is required because no controller-processor relationship exists. The tool processes data entirely in your browser memory, which GDPR treats as within your own processing environment rather than a third-party processing event. Consequently, the scrub step adds zero GDPR compliance overhead to your workflow. The scrubbed output, containing only tokens rather than personal data, may be sent to an AI provider without triggering Article 46 transfer requirements, because tokens are not personal data under GDPR and carry no information that can identify an individual even if the AI provider stores the prompt indefinitely.
GDPR Article 35 Data Protection Impact Assessments and AI tool use
GDPR Article 35 requires a Data Protection Impact Assessment before any processing likely to result in high risk to the rights and freedoms of natural persons. Using an AI tool to process personal data at scale qualifies as likely-high-risk processing in most DPA guidance, because the processing involves systematic evaluation or profiling of individuals, uses new technology, or occurs at significant scale. Organizations that routinely use AI tools on customer or employee data without a completed DPIA expose themselves to enforcement under Article 83(4), which carries fines of up to €10 million or 2% of global annual turnover5.
A DPIA for AI tool use must describe the processing purposes, assess necessity and proportionality, evaluate risks to data subjects, and identify mitigating measures. List pre-submission tokenization as a DPIA safeguard: it is a documented technical measure that directly mitigates the risk of unauthorized disclosure to the AI provider, and naming it explicitly gives assessors a concrete control to point to in the DPIA documentation. Several European DPAs, including the French CNIL and the UK ICO, publish template DPIAs for AI tool use that you can adapt for your organization's specific context.
DPIA checklist for teams using AI for personal data processing
A practical DPIA checklist for teams using AI tools with personal data includes: identifying the Article 6 legal basis for general personal data processing (and the Article 9 condition for special categories), confirming a valid transfer mechanism exists if the AI provider is outside the EU (SCCs under Article 46 or an adequacy decision), documenting the categories of personal data involved, assessing whether data minimization measures including pre-submission scrubbing are in place, and obtaining a written opinion from your Data Protection Officer before processing begins. Each of these checklist items has a direct corresponding GDPR obligation, making the DPIA both a compliance tool and an operational planning document.
Schrems II and transatlantic AI data transfers
The Schrems II judgment (Data Protection Commissioner v. Facebook Ireland, C-311/18) invalidated the EU-US Privacy Shield in 20206 and established that Standard Contractual Clauses alone are insufficient when the recipient country's surveillance law makes effective data protection impossible. US cloud providers including Google, Microsoft, and OpenAI are subject to FISA Section 702 and Executive Order 123337, which permit US government access to data stored on US infrastructure. For EU organizations sending personal data to US-based AI providers, the legal basis for the transfer requires either SCCs supplemented by additional technical safeguards, or BCRs reviewed by a competent DPA.
The EU-US Data Privacy Framework (adopted July 2023) established a new adequacy mechanism for transfers to certified US companies8. US AI providers certified under the DPF can receive EU personal data under an adequacy decision, which simplifies the legal basis compared to the SCC route. Verify DPF certification on the official DPF Program Website before relying on it as the transfer mechanism for your AI tool data flows. Certification is voluntary and must be renewed annually by the US company, so an organization certified today may become uncertified if it fails to renew.
Selecting EU-hosted AI options for maximum GDPR coverage
Several AI providers offer EU-hosted deployment options that keep data within the European Economic Area. Microsoft Azure OpenAI Service can be deployed in the West Europe or North Europe Azure regions. Google Cloud Vertex AI supports EU-only data residency for Organization customers. Mistral AI, headquartered in France, operates EU infrastructure for its La Plateforme API. Choosing an EU-hosted provider eliminates the transatlantic transfer concern entirely and removes the need to verify SCCs or DPF certification for the AI processing step, simplifying both the DPIA and the legal basis documentation for EU teams.
Scrubbing before submission complements an EU-hosted provider by removing identifiers regardless of where the AI processes the data. CapyToolkit runs locally in your browser, so no personal data reaches any server during scrubbing, and the tokens that result are not personal data under GDPR, which keeps the cross-border transfer question closed even for non-EU providers.
When to use this
Use this before sending any prompt containing EU personal data to an AI provider, especially when no valid DPA exists with that provider or when cross-border transfer mechanisms are unclear.
Examples
Customer support prompt with EU customer details
Help me draft a refund email to Lena Müller ([email protected], +49-89-123456) about order #DE-2026-00123.
Help me draft a refund email to Lena Müller ([EMAIL_1], [PHONE_1]) about order #DE-2026-00123.
The name is retained here. Email and phone — direct personal data identifiers under GDPR — are tokenized.
GDPR data subject request analysis
We received a DSAR from [email protected] (IP: 85.214.132.117). List which data fields we hold.
We received a DSAR from [EMAIL_1] (IP: [IP_1]). List which data fields we hold.
- 1.
DLA Piper, "GDPR Fines and Data Breach Survey: January 2026," dlapiper.com, January 2026. https://www.dlapiper.com/insights/publications/2026/01/dla-piper-gdpr-fines-and-data-breach-survey-january-2026
- 2.
CJEU, "Judgment of the Court (Second Chamber) of 19 October 2016, Patrick Breyer v Bundesrepublik Deutschland (C-582/14)," eur-lex.europa.eu, October 2016. https://eur-lex.europa.eu/legal-content/en/TXT/?uri=CELEX%3A62014CJ0582
- 3.
GDPR, "Art. 5 – Principles relating to processing of personal data," gdpr-info.eu, accessed June 2026. https://gdpr-info.eu/art-5-gdpr/
- 4.
GDPR, "Art. 28 – Processor," gdpr-info.eu, accessed June 2026. https://gdpr-info.eu/art-28-gdpr/
- 5.
UK Legislation, "Regulation (EU) 2016/679 – Article 83," legislation.gov.uk, accessed June 2026. https://www.legislation.gov.uk/eur/2016/679/article/83
- 6.
"Schrems II," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/Schrems_II
- 7.
"Executive Order 12333," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/Executive_Order_12333
- 8.
European Commission, "Standard Contractual Clauses (SCC)," commission.europa.eu, accessed June 2026. https://commission.europa.eu/law/law-topic/data-protection/international-dimension-data-protection/standard-contractual-clauses-scc_en
Yes, in most contexts. The CJEU confirmed in Breyer v. Germany (2016) that dynamic IP addresses can be personal data. The scrubber detects and tokenizes both IPv4 and IPv6 addresses.
If the data sent to the AI provider contains no personal data (only tokens), there is no personal data transfer to govern under Article 46. Consult your DPO for a formal determination, as GDPR compliance depends on the specific processing context.
No. All processing occurs in your browser, and no data is transmitted to any server. Without a controller-processor relationship, no DPA is needed for the scrubbing step.
Names are not detected by pattern matching since they are too variable to reliably distinguish from other text. For prompts that include full names as personal data, remove or replace names manually before using the scrubber for the remaining fields.
Yes. Even with a GDPR-compliant AI provider, scrubbing is a good data minimization practice. The scrubber reduces what you send to the AI to what is strictly necessary, which is a principle GDPR endorses regardless of the provider's certification status.
CCPA Data Scrubber: Protect California Consumer Data
CCPA covers any business that handles California residents' data above its revenue or volume thresholds1, which includes most mid-size companies operating in the US. Sharing California consumer data with AI tools without a service provider agreement constitutes a "sale" or "share" under CCPA, triggering opt-out rights and potential fines from the California Privacy Protection Agency.
The California Privacy Rights Act amendments added the right to limit use of sensitive personal information, including SSNs, account log-in credentials, and health data2. Scrubbing these fields before AI submission satisfies the limitation right technically: the sensitive data never reaches the AI provider, so its use there is not a concern.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →CCPA categories the scrubber addresses
CCPA defines specific categories of personal information. The scrubber directly targets: identifiers (email addresses, phone numbers, IP addresses), financial information (credit card numbers, IBAN bank account numbers), internet or network activity (IP addresses in web logs, JWT session tokens), and sensitive personal information (SSNs, account log-in information in the form of API keys and database credentials). Building on this, CCPA's "household" concept means that an IP address associated with a home network is personal information for any California residents at that address3, and IPv4 and IPv6 detection covers this category.
The CCPA category that most frequently triggers AI-related violations is identifiers, because email addresses and IP addresses appear in virtually every log file, support ticket, and config export that teams paste into AI tools. A single customer support log pasted for AI-assisted analysis can contain the customer's email, their IP address from the session record, and an internal API key from the error message, all of which qualify as personal information under CCPA and must be removed before the data reaches an uncertified AI provider.
Service provider agreements and AI tools
Under CCPA, sharing personal information with an AI provider without a qualifying service provider agreement constitutes a "sale" or "share"4, even when no money changes hands. OpenAI, Anthropic, Google, and Microsoft offer data processing agreements for their enterprise tiers5, but most teams using free or standard API tiers have no qualifying agreement in place. Consequently, sending California consumer data to a standard AI API tier may violate CCPA regardless of the provider's privacy reputation. Scrubbing before transmission removes the personal information, so the shared content does not qualify as personal information under CCPA.
The practical implication for California businesses is significant: every employee who pastes customer data into a free-tier AI tool without a service provider agreement in place creates a potential CCPA violation, and the California Privacy Protection Agency has increased enforcement focus on exactly this kind of AI-related data sharing in 2026. Pre-scrubbing the personal identifiers before the prompt reaches the AI provider eliminates the violation at the source, because the AI provider receives only tokens that do not qualify as personal information under CCPA's definition.
CPRA sensitive personal information
The California Privacy Rights Act added a new category: sensitive personal information. This includes SSNs, driver's license numbers (not yet detected), account log-in credentials, precise geolocation, racial or ethnic origin, and email content (email address detected). Yet the scrubber covers the most commonly digitized sensitive fields, including SSNs, email addresses, and account credentials, which are the categories most frequently embedded in operational text pasted into AI tools.
For regulated documents containing driver's license or passport numbers, manual review is also needed because these identifiers do not follow patterns that regex can reliably detect. The CPRA also gives California consumers the right to limit how businesses use their sensitive personal information6, which means that even after scrubbing, organizations should document which data categories were removed and retain that documentation for compliance purposes.
CPRA's right to limit sensitive personal information and what it requires
The California Privacy Rights Act, effective January 1, 2023, added a right for consumers to limit how businesses use and disclose their sensitive personal information. Sensitive personal information under CPRA (Section 1798.121) includes SSNs, driver's license numbers, account log-in credentials combined with security codes, precise geolocation, racial or ethnic origin, religious beliefs, union membership, mail contents, genetic data, biometric information used for identification, and information about a consumer's sex life or sexual orientation.
When a consumer exercises the right to limit, businesses must stop using that consumer's sensitive personal information for purposes beyond providing the requested goods or services. Using a California consumer's SSN, account credentials, or precise location data as input to an AI tool for purposes unrelated to the service contract (such as internal analytics, product development, or model training) likely exceeds the permitted use under CPRA. Scrubbing sensitive personal information from prompts before AI submission ensures compliance with limit-use elections regardless of whether any specific consumer has yet made such an election.
Responding to a right-to-limit request under CPRA
When a California consumer submits a right-to-limit request for their sensitive personal information, your response must identify all processing of that consumer's SPI and stop using it for unauthorized purposes within 15 business days7. For AI workflows that have involved the consumer's SPI in previous prompts, you cannot remove that data from an external AI provider's systems retroactively. Scrubbing SPI before submission prevents this retroactive problem from arising: if the AI provider never received the SPI, no retroactive deletion request is needed and the right-to-limit obligation is satisfied by the technical control rather than by vendor negotiation.
CCPA service provider agreements and AI tool requirements
Under CCPA Section 1798.100(d), sharing personal information with a service provider requires a written contract prohibiting the service provider from retaining, using, or disclosing personal information for any purpose other than the specified business purpose8. Most standard-tier AI API agreements do not satisfy this requirement. OpenAI's standard API terms, Google's Gemini API terms, and Anthropic's standard Claude API terms permit the provider to use submitted data for model improvement by default. Only enterprise agreements with explicit prohibition clauses and zero-data-retention commitments qualify as service provider agreements under CCPA.
Businesses using AI APIs without a qualifying service provider agreement risk treating the data transmission as a "sale" or "share" under CCPA, which triggers consumer opt-out rights under Section 1798.120 and potential enforcement by the California Privacy Protection Agency. The CPPA increased its enforcement focus on AI-related data practices in 20269. Scrubbing California consumer personal information before AI submission eliminates the sale-or-share classification for the AI interaction, because the AI provider receives only tokens that contain no personal information as defined by CCPA.
Identifying which AI provider agreements qualify as CCPA service provider agreements
Review your AI provider's Data Processing Addendum before classifying an AI tool as a CCPA service provider. The qualifying DPA must include: a prohibition on retaining personal information beyond the service period, a prohibition on using personal information for any purpose other than performing the contracted services, a prohibition on combining personal information from your business with personal information from other sources, and a certification that the provider understands and will comply with CCPA. OpenAI's Enterprise Data Processing Agreement, Google's Cloud Data Processing Addendum (for Workspace and Cloud), and Anthropic's Enterprise Privacy Agreement each address these requirements for their respective enterprise tiers.
California consumer data in marketing AI workflows
Marketing teams use AI tools heavily for campaign copy generation, audience segmentation analysis, and customer communication drafting. These tasks frequently involve California consumer data in the form of email campaign lists, contact engagement histories, and CRM segments. Uploading a campaign list to an AI tool for segmentation analysis, or using customer engagement data as prompt context for copy personalization, constitutes sharing personal information with a third-party AI provider for a marketing purpose, which CCPA Section 1798.120 covers as a consumer opt-out right.
For marketing workflows involving California consumers, apply the scrubber to email addresses, phone numbers, and any other personal identifiers in the prompt before submitting to the AI. An AI writing assistant can generate personalized campaign copy from a tokenized customer profile (customer segment, purchase history category, geographic region without street address) just as effectively as from a record with the real email and phone. Replace customer identifiers with segment descriptors before the prompt ever reaches the AI.
Filtering opted-out California consumers before AI-assisted campaigns
California consumers who have submitted opt-out requests under CCPA Section 1798.120 must be excluded from data sharing with third parties for marketing purposes. Before using a contact list in any AI tool for campaign purposes, filter out opted-out California consumers based on their declared California residence or IP geolocation. The remaining list may still require scrubbing of personal identifiers for CCPA compliance with AI providers that lack qualifying service provider agreements. Maintaining a current opt-out list in your CRM and applying it as a filter at the list export step, before scrubbing, ensures both the opt-out and data-minimization obligations are satisfied simultaneously.
Filtering and scrubbing together satisfy both the opt-out and data-minimization obligations before the prompt leaves your system. Because CapyToolkit processes the list locally in your browser, no California consumer record is uploaded to a server during scrubbing, and the tokens that result are not personal information under CCPA, which keeps the AI interaction outside the sale-or-share definition.
When to use this
Use this before sending any prompt about California residents, including customers, employees, or prospects, to an AI tool when a qualifying service provider agreement with that AI provider is not in place.
Examples
Customer service query with California resident data
Help draft a response to Maria Garcia ([email protected], SSN 321-54-9876) who filed a DSAR under CCPA.
Help draft a response to Maria Garcia ([EMAIL_1], SSN [SSN_1]) who filed a DSAR under CCPA.
The CCPA context is preserved for the AI. The identifiers that are personal information under CCPA are tokenized.
Marketing data analysis with IP addresses
Analyze engagement patterns: user 203.0.113.5 clicked 3 times, user 198.51.100.2 bounced immediately.
Analyze engagement patterns: user [IP_1] clicked 3 times, user [IP_2] bounced immediately.
- 1.
CPPA, "Updated Monetary Thresholds in CCPA," cppa.ca.gov, effective January 1, 2025. https://www.cppa.ca.gov/regulations/cpi_adjustment.html
- 2.
CPPA, "CalPrivacy Brings New Round of Enforcement Actions Against Data Brokers," cppa.ca.gov, January 2026. https://www.cppa.ca.gov/announcements/2026/20260108.html
- 3.
California Legislature, "Civil Code § 1798.140 – Definitions," leginfo.legislature.ca.gov, accessed June 2026. https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- 4.
IAPP, "Analyzing the CPRA's New Contractual Requirements for Transfers of Personal Information," iapp.org, March 2021. https://iapp.org/news/a/analyzing-the-cpras-new-contractual-requirements-for-transfers-of-personal-information
- 5.
AI Policy Desk, "Privacy-First AI APIs: Which Don't Train on Your Data in 2026," aipolicydesk.com, April 2026. https://www.aipolicydesk.com/blog/privacy-first-ai-api-no-training-gdpr-ccpa-2026
- 6.
California Legislature, "Civil Code § 1798.121 – Right to Limit Use and Disclosure of Sensitive Personal Information," leginfo.legislature.ca.gov, accessed June 2026. https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.121
- 7.
Cornell LII, "Cal. Code Regs. Tit. 11, § 7027 – Requests to Limit Use and Disclosure of Sensitive Personal Information," law.cornell.edu, accessed June 2026. https://www.law.cornell.edu/regulations/california/11-CCR-7027
- 8.
Cornell LII, "Cal. Code Regs. Tit. 11, § 7051 – Service Providers and Contractors," law.cornell.edu, accessed June 2026. https://www.law.cornell.edu/regulations/california/11-CCR-7051
- 9.
OpenAI, "Data Processing Addendum," openai.com, February 2024. https://openai.com/policies/feb-2024-data-processing-addendum/
CCPA defines 'sale' broadly to include sharing personal information for value. Even without payment, sharing consumer data with an AI provider that uses it for model improvement or analytics may qualify. A qualifying service provider agreement that restricts the recipient's use of the data is needed to avoid this classification.
Yes. CPRA added sensitive personal information rights, stronger opt-out requirements, and created the California Privacy Protection Agency with independent enforcement power. The scrubber addresses the sensitive personal information category by removing SSNs, account credentials, and email addresses.
The right to delete is a separate obligation that requires deleting consumer data from your records upon request. Scrubbing before AI submission addresses a different risk by preventing unauthorized disclosure to third parties. These are complementary, not interchangeable, obligations.
Physical addresses, purchase and browsing history patterns, inferences drawn from data, and biometric data are not covered by the current detection set. These categories require manual review for CCPA-sensitive use cases.
No. The scrubber processes text locally in your browser, and no personal information reaches any external server. Without a controller-processor relationship, no service provider agreement is needed or available.
HIPAA PHI Scrubber: Remove Patient Data Before AI
PHI violations cost healthcare organizations millions. The average HIPAA breach fine exceeds $1.2 million1, and OCR enforcement in 2026 has focused specifically on organizations that allowed PHI to reach AI tools without a Business Associate Agreement2. A clinician pasting patient notes into ChatGPT for summarization is a HIPAA violation, unless the PHI is removed before the notes reach the AI provider's servers.
The scrubber removes identifiable health information fields before they leave your device. Because CapyToolkit is a local browser tool that never receives PHI, no BAA with CapyToolkit is needed, and the scrubbing step is HIPAA-neutral. Only the scrubbed, de-identified output reaches the AI provider, which then does not require a BAA for that specific content.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →The 18 HIPAA Safe Harbor identifiers
HIPAA de-identification method requires removing 18 specific identifiers from patient data before it can be shared freely3. The scrubber detects and tokenizes 8 of these directly using pattern matching: email addresses, phone numbers including fax, Social Security Numbers, account numbers detected via IBAN patterns, IP addresses, device identifiers caught through JWT and credential detection, geographic subdivisions identified via internal domain detection, and certain health-system contact information embedded in clinical text. Furthermore, the scrubber catches credit card numbers validated by the Luhn algorithm and authentication tokens that appear in healthcare administrative data such as billing records and insurance forms. The remaining 10 HIPAA identifiers, including patient names, detailed geographic data, biometric identifiers, and most dates associated with care events, require manual review and removal since these do not follow predictable text patterns that regex can reliably detect.
Healthcare data types the scrubber covers
In healthcare workflows, the most commonly AI-assisted tasks involve clinical documentation, billing queries, and patient communication drafting. Clinical notes pasted for AI summarization often contain patient email addresses, phone numbers, and embedded SSNs from intake forms. Billing data includes insurance member IDs, credit card information for copay collection, and IBAN references for international healthcare providers. Building on this, staff emails and internal clinical system hostnames appear in IT tickets and helpdesk queries. The scrubber detects all these patterns, reducing the identifiable content before it reaches any AI provider.
Prior authorization documents deserve special attention because they combine diagnosis codes, medication names, and clinical justification narratives with the patient contact fields that make the document PHI. When a revenue cycle team uses an AI tool to draft or review a prior authorization letter, the scrubber removes the email, phone, and SSN fields while preserving the clinical codes and procedure descriptions that the AI needs to generate an accurate narrative. This selective removal keeps the document useful for AI assistance without transmitting the identifiers that create HIPAA exposure.
BAA requirements and local tools
A Business Associate Agreement is required with any vendor that creates, receives, maintains, or transmits PHI on behalf of a covered entity. CapyToolkit's scrubber does not receive PHI, as it processes text locally in your browser memory, and consequently no BAA with CapyToolkit is required for the scrubbing step. The AI provider you send scrubbed text to may still require a BAA if any residual PHI remains in the scrubbed output, so scrubbing before transmission reduces residual PHI exposure but does not guarantee complete de-identification under HIPAA's expert determination method.
How the scrubber avoids the BAA requirement
A cloud-based PII scrubbing service that receives your text on its servers is itself a business associate under HIPAA and requires a BAA before any PHI reaches its infrastructure. Understanding this distinction matters when healthcare organizations evaluate their AI tool chain and choose where to draw the BAA boundary, because the browser-based scrubber avoids the BAA requirement entirely by ensuring that PHI processing happens only within the clinician's browser tab. This architectural difference has direct implications for how covered entities manage their vendor risk and BAA negotiation workflows.
The 18 HIPAA Safe Harbor identifiers and which the scrubber addresses
HIPAA's Safe Harbor de-identification standard requires removing all 18 identifier categories before health data can be shared freely without a BAA3. The 18 categories are names, geographic subdivisions smaller than a state, all dates except year, phone numbers, fax numbers, email addresses, Social Security Numbers, medical record numbers, health plan beneficiary numbers, account numbers, certificate and license numbers, vehicle identifiers, device identifiers, web URLs, IP addresses, biometric identifiers (fingerprints, voiceprints), full-face photographs, and any other unique identifying number or code. The scrubber directly addresses 7 of these 18 through pattern detection: phone numbers, fax numbers, email addresses, Social Security Numbers, account numbers via IBAN detection, IP addresses, and device identifiers via JWT and API key detection.
The remaining 11 identifier categories, including names, geographic data below state level, dates, medical record numbers, health plan beneficiary numbers, certificate numbers, vehicle identifiers, URLs, biometric identifiers, and photographs, require manual review and removal before the scrubbed text meets Safe Harbor criteria. Meeting the full standard requires both the automated scrubbing pass and a careful human review of the output for these remaining categories. For clinical documents requiring Safe Harbor de-identification, use the scrubber first to remove the pattern-detectable fields, then manually review the output for names, dates, and geographic information.
Manually removing identifiers outside the scrubber's detection scope
A practical workflow is to layer manual redaction on top of the automated pass: replace patient names with initials, replace specific dates with the year only (admitting year 2026 instead of the exact admission date), and replace ZIP codes with the first three digits only. Three-digit ZIP codes represent populations large enough for de-identification in most US counties. After both the automated and manual steps, review the output against the 18-category checklist to confirm all identifiers have been addressed before sharing with any party that does not hold a BAA.
AI platforms with signed BAAs for healthcare workflows
Several AI providers offer signed Business Associate Agreements for healthcare organizations. Microsoft Azure Health AI services (including Azure OpenAI Service for Healthcare) provide a standard BAA through the Microsoft Online Subscription Agreement when Azure is purchased for healthcare use. AWS HealthLake and Amazon Comprehend Medical are HIPAA-eligible services covered under AWS's standard BAA for healthcare customers4.
Google Cloud Healthcare API and generative AI services in Vertex AI are available under Google Cloud's BAA for qualifying healthcare customers. General-purpose consumer AI tools, including ChatGPT free and Plus tiers, Claude.ai without an enterprise agreement, and Gemini standard accounts, do not offer BAAs and are not HIPAA-eligible. Healthcare organizations that need AI assistance for clinical tasks should either use a BAA-covered platform or maintain a strict scrubbing workflow that removes all 18 Safe Harbor identifiers.
Verifying HIPAA eligibility for a specific AI service
AWS publishes its list of HIPAA-eligible services at aws.amazon.com/compliance/hipaa-eligible-services-reference. Microsoft publishes an equivalent list at microsoft.com/en-us/trustcenter/compliance/hipaa. Not all services from a HIPAA-eligible provider are themselves covered, so an AWS customer with a BAA must verify that the specific service (for example, Amazon Bedrock versus Amazon Comprehend Medical) is on the eligible services list before using it to process PHI. Relying on a provider's general HIPAA eligibility without service-level verification is a common compliance gap that OCR investigations have identified.
Verifying the service tier and scrubbing before submission are complementary controls, because eligibility alone does not remove identifiers that remain in the prompt. CapyToolkit runs the scrubber locally in your browser with no BAA needed, so the PHI is tokenized before it reaches any provider, and only the variables file maps the tokens back to the real patient values.
Minimum Necessary standard in healthcare AI prompts
HIPAA's Minimum Necessary standard (45 CFR 164.502(b)) requires covered entities to use, disclose, and request only the minimum PHI necessary to accomplish the intended purpose5. Applying this standard to AI prompt design means including only the specific identifiers and clinical details the AI needs for the task, not the full patient record. A prompt asking an AI to generate a prior authorization letter for a specific medication needs the patient's insurance plan type, the diagnosis code, and the medication name. It does not need the patient's phone number, SSN, or full medical history.
Designing minimal prompts before adding the scrubber layer eliminates the most sensitive PHI before the scrubber even runs. Structure clinical AI prompts as role-based templates that include categories of information rather than individual patient records. For example, write "Patient is a 65-year-old female with a history of Type 2 diabetes, requesting prior authorization for medication X under insurance plan Y" instead of including the patient's real name and specific dates.
Replace specific ages with ranges, replace specific dates with relative references such as three months ago or during the prior calendar year, and omit names entirely when the clinical task does not require them. Maintain a library of approved AI prompt templates that your Privacy Officer has reviewed against the Minimum Necessary standard, documenting which templates were used in AI-assisted tasks to create an audit trail demonstrating deliberate compliance.
When to use this
Use this before pasting any clinical documentation, patient communication, billing record, or healthcare administrative text into an AI tool when the content contains identifiable patient information.
Examples
Clinical note for AI-assisted summarization
Patient: M.T., DOB 1975-03-14, SSN: 321-54-9876. Complaint: chest pain. Contact: [email protected], (555) 321-9876.
Patient: M.T., DOB 1975-03-14, SSN: [SSN_1]. Complaint: chest pain. Contact: [EMAIL_1], [PHONE_1].
DOB and initials are preserved. SSN, email, and phone are the most directly identifying fields and are all tokenized.
Billing query with insurance and payment details
Claim for patient [email protected], insurance ID: INS-123456, card: 5500000000000004.
Claim for patient [EMAIL_1], insurance ID: INS-123456, card: [CC_1].
- 1.
HIPAA Journal, "HIPAA Violation Fines," hipaajournal.com, updated June 2026. https://www.hipaajournal.com/hipaa-violation-fines/
- 2.
HIPAA Journal, "2026 HIPAA Violation Fines and Settlements," hipaajournal.com, accessed June 2026. https://www.hipaajournal.com/2025-healthcare-data-breach-report/
- 3.
HHS OCR, "Guidance Regarding Methods for De-identification of Protected Health Information," hhs.gov, accessed June 2026. https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html
- 4.
Microsoft, "Health Insurance Portability and Accountability Act (HIPAA) & HITECH Act," learn.microsoft.com, accessed June 2026. https://learn.microsoft.com/en-us/compliance/regulatory/offering-hipaa-hitech
- 5.
HHS OCR, "Minimum Necessary Requirement," hhs.gov, accessed June 2026. https://www.hhs.gov/hipaa/for-professionals/privacy/guidance/minimum-necessary-requirement/index.html
The scrubber removes the most commonly digitized HIPAA Safe Harbor identifiers it can detect by pattern. It does not address all 18 Safe Harbor identifiers, and names, geographic data, and certain dates require manual removal. The scrubber is a practical risk reduction tool, not a certified de-identification system.
No. CapyToolkit is a local browser tool that never receives, stores, or transmits PHI. The processing happens in your browser memory. No controller-processor relationship exists, so no BAA is required.
If residual PHI remains in the scrubbed output, yes. The scrubber removes identifiable patterns it detects, so if additional PHI (names, dates, geographic data) is present, the AI provider still receives that PHI. For healthcare organizations, a BAA with the AI provider remains recommended even when using the scrubber.
Patient names, geographic data (zip codes, cities, states), dates associated with the patient, ages, full face photographs, and biometric identifiers are not detected by the current pattern set. These require manual removal.
The scrubber is a browser-based tool that can be accessed from any network. Hospital IT policies govern what URLs staff may access. The tool itself processes data locally and makes no external network calls during operation.
Scrub PHI and PII from Healthcare Data Before AI
Healthcare data holds the most sensitive personal information a person can share. Patient records combine direct identifiers such as name, SSN, and date of birth with health status information, diagnoses, medications, and procedures in a single document. When healthcare professionals use AI tools to draft clinical documentation, analyze case histories, or generate patient communications, any of these fields in the prompt creates a HIPAA exposure.
The scrubber removes the identifiable fields that make health data protected, the direct identifiers, before the document reaches an AI provider. Clinical content and medical context remain intact, giving the AI model what it needs to assist without receiving what HIPAA protects.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →Protected health information in clinical workflows
HIPAA protects 18 categories of identifiers when combined with health information1. The most frequently digitized identifiers in clinical workflows are patient email addresses from intake forms, phone numbers from contact records, Social Security Numbers from insurance claims, and IP addresses from patient portal access logs. Furthermore, clinical system hostnames such as .hospital.org and .clinic.internal used in IT support queries are organizational identifiers that link the system to specific patients. The scrubber detects all of these pattern types and reduces the identifiable surface of a clinical document before AI processing.
The pattern-detectable identifiers are only part of the picture, however. A clinical document that has been scrubbed of emails, phone numbers, and SSNs may still contain patient names, admission dates, and geographic subdivisions that qualify as PHI under HIPAA. The scrubber addresses the machine-readable fields efficiently, but a complete de-identification workflow requires a second pass for the identifiers that regex cannot catch, particularly patient names embedded in clinical narrative text.
Clinical documentation use cases
AI is increasingly used for clinical documentation assistance: transcribing patient-physician conversations, summarizing case histories, generating discharge summaries, and drafting referral letters. Each of these documents is rich in PHI. Building on this, administrative tasks such as billing query resolution, insurance preauthorization letters, and patient complaint responses also touch patient identifiers. The scrubber is most practically used at the copy-paste step: before the clinician copies a section of the record into the AI tool, they paste it through the scrubber first. The scrubbed text retains the clinical meaning while removing the identifying markers.
Referral letters are a high-value target for scrubbing because they combine the patient's clinical history with their contact information, insurance details, and the referring provider's identity, all in a single document. When a specialist pastes a referral letter into an AI tool to help draft a response or summarize the clinical question, the scrubber removes the patient email, phone, and SSN while preserving the diagnosis codes, medication lists, and clinical narrative that the AI needs to generate a useful response.
Limitations for full HIPAA compliance
The scrubber is a practical risk reduction tool, not a certified HIPAA de-identification system. It addresses pattern-detectable identifiers but cannot detect patient names, geographic sub-unit data, dates associated with the patient, face photographs, or biometric data. Consequently, clinical documents run through the scrubber still require manual review for names and dates before being passed to an AI tool in a strictly HIPAA-compliant workflow. The scrubber handles the machine-detectable majority, and human review handles the remainder.
For healthcare organizations that require certified de-identification, the HIPAA expert determination method provides a formally recognized alternative to Safe Harbor2. Under this method, a qualified statistician certifies that the risk of re-identification is very small, even if some identifiers remain. The scrubber can serve as the first step in an expert-determination workflow: the automated pass removes the easily detectable fields, and the human expert evaluates the residual risk from the remaining quasi-identifiers such as rare diagnoses, unusual age values, or small-sample geographic data.
EHR systems and the clinical copy-paste intervention point
Electronic health record systems (Epic, Cerner, Oracle Health, Meditech) display patient data as a combination of structured field values and free-text clinical narrative. Clinicians copy sections of EHR notes for AI-assisted documentation, discharge summary generation, and referral letter drafting. The copy-paste step between the EHR and the AI tool is the correct intervention point for scrubbing: the scrubber runs on the clipboard content before it is pasted into the AI interface, with no EHR system integration or API access required.
Paste the copied EHR section into the scrubber, copy the scrubbed version, then paste into the AI tool. The AI receives the clinical narrative structure with identifiable fields replaced by tokens. For discharge summary drafting, the AI produces a structured summary from the scrubbed clinical context; the care team then reviews the output and reinserts patient-specific values as needed before finalizing3. Oracle Health (formerly Cerner) embeds its Clinical AI Agent directly into EHR workflows, demonstrating that AI-assisted documentation is becoming native to the systems clinicians already use4. This workflow requires no changes to EHR configuration and no vendor relationship with any AI provider on the healthcare organization's behalf.
Integration with EHR ambient documentation workflows
AI ambient documentation tools (clinical note generation from conversation audio) present a different scrubbing challenge: the text is generated by the AI rather than pasted by the clinician. Review the output from ambient documentation tools for HIPAA identifiers before including it in any subsequent AI processing step. The scrubber applies to any text: ambient documentation output can be pasted into the scrubber to remove identifiers before the note is sent to a second AI tool for formatting or summarization.
Billing and prior authorization workflows with PHI and financial data
Insurance prior authorization letters require linking patient identifiers with diagnosis codes, procedure codes, and clinical justification narratives. Revenue cycle management teams use AI tools to draft denial appeals, write prior authorization letters, and summarize clinical justifications for payors. These documents combine financial identifiers (member ID, group number, insurance plan name) with protected health information, qualifying as PHI under HIPAA in a healthcare organizational context.
Scrubbing removes the contact and identity fields that make the document PHI while preserving the clinical codes, procedure descriptions, and financial logic that the AI needs to draft or review the document. An AI that sees [EMAIL_1] for the patient contact and [PHONE_1] for the billing phone can still produce a correctly structured prior authorization narrative. The payor name and insurance plan information do not qualify as patient PII under HIPAA and can remain in the prompt without scrubbing.
Remittance advice and explanation-of-benefits processing
Remittance advice documents (ERA files, 835 EDI transactions) and explanation-of-benefits statements from payors contain patient identifiers, service dates, claim amounts, and adjustment reason codes. AI tools are used to analyze these documents for billing error identification and denial pattern recognition. Pull member IDs, keep the adjustment codes when preparing ERA text for AI analysis; the claim amounts, procedure codes, and adjustment reason codes that reveal billing patterns are not PHI and do not require removal.
Patient communication templates and AI-assisted personalization
AI tools draft appointment reminders, care gap outreach letters, discharge instructions, and post-visit follow-up messages for healthcare organizations. Personalizing these communications requires patient name, phone number, preferred email, appointment date, and sometimes insurance information. Each of these fields is a HIPAA identifier when combined with the health context of the communication.
Use the scrubber to remove personalization identifiers before AI drafting, then have the care team reinsert the correct patient-specific values in the final document. The AI produces the structural template (the language for a missed appointment reminder, the explanation for a care gap notification) from the scrubbed context. Staff then replace [EMAIL_1], [PHONE_1], and [NAME_placeholder] (manually added before scrubbing for name fields) with the actual patient values in the final step. This workflow separates the AI's structural work from the patient-specific data, keeping PHI within the healthcare organization's controlled environment throughout the drafting process.
Secure messaging platforms and AI integration risk
Healthcare organizations using secure messaging platforms (TigerConnect, Klara, Spruce Health) sometimes enable AI features within those platforms for message summarization or response drafting. These platform-embedded AI features transmit message content to the platform's AI provider. TigerConnect, for example, uses role-aware AI agents that integrate data from EHRs and route the right information to the right team member in real time, which means message content may pass through AI processing within the platform itself5. Review the BAA status of your secure messaging platform's AI features before enabling them, and apply the same scrubbing discipline to message content shared across AI-enabled channels within the platform.
Treating message content with the same scrubbing discipline as any other prompt keeps patient identifiers inside the platform's protected environment. CapyToolkit runs the scrubber locally in your browser, so the clinical text never leaves the device during processing, and the variables file you download is the only record that maps the tokens back to the real PHI.
When to use this
Use this before pasting clinical notes, patient communications, billing records, or healthcare administrative documents into any AI tool when the document contains identifiable patient information.
Examples
Discharge summary for AI-assisted generation
Patient: R.J. Martinez, DOB 1968-07-22, SSN 456-78-9012. Admitted 2026-05-15. Dx: T2DM. Attending: [email protected].
Patient: R.J. Martinez, DOB 1968-07-22, SSN [SSN_1]. Admitted 2026-05-15. Dx: T2DM. Attending: [EMAIL_1].
Patient portal message with contact details
Message from [email protected] (Cell: (408) 555-0199): My prescription wasn't ready. Account #: 1234-5678.
Message from [EMAIL_1] (Cell: [PHONE_1]): My prescription wasn't ready. Account #: 1234-5678.
- 1.
HHS OCR, "De-identification of Protected Health Information," hhs.gov, accessed June 2026. https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html
- 2.
Accountable HQ, "HIPAA De-Identification Explained: Safe Harbor, Expert Determination, and Risk Controls," accountablehq.com, accessed June 2026. https://www.accountablehq.com/post/hipaa-de-identification-explained-safe-harbor-expert-determination-and-risk-controls
- 3.
Epic, "AI Charting," epic.com, February 2026. https://www.epic.com/epic/post/ai-charting/
- 4.
Oracle Health, "Oracle Health — Clinical AI," oracle.com, accessed June 2026. https://www.oracle.com/health/
- 5.
TigerConnect, "AI That Orchestrates Care," tigerconnect.com, accessed October 2026. https://tigerconnect.com/resources/videos/ai-that-orchestrates-care-lp/
No. The scrubber covers approximately 8 of the 18 by pattern: email addresses, phone numbers, SSNs, account numbers (via IBAN), IP addresses, device identifiers (as JWT tokens), and clinical system hostnames. The remaining 10, including names, geographic subdivisions, dates, ages, and biometric identifiers, require manual removal.
No. Clinical EHR de-identification tools are purpose-built for HIPAA compliance and use expert determination methods. The scrubber is a practical, zero-setup tool for reducing PHI exposure at the prompt level, not a replacement for certified de-identification infrastructure.
No BAA is available or needed. CapyToolkit is a local browser tool that does not receive, transmit, or process PHI. The scrubbing happens in your browser memory only.
Mental health records are subject to additional protections beyond standard HIPAA under 42 CFR Part 2 for substance use disorder records and various state laws. The scrubber addresses the pattern-detectable identifiers in these records but does not provide specialized protections for this category.
Yes. The tool requires no IT setup, no software installation, and no account. Any staff member with a browser can access and use it immediately.
PCI DSS Data Masking: Remove Cardholder Data Before AI
Cardholder data exposed to AI is a PCI DSS violation. PAN numbers, cardholder names combined with expiry dates, CVV codes, and card validation data are all in scope for PCI DSS, and sending them to any AI tool, even for fraud analysis or customer dispute review, is a breach of Requirement 4 unless that tool is within your PCI DSS scope1.
Masking cardholder data before it reaches an AI tool keeps the AI provider out of your PCI DSS scope entirely. The scrubber detects credit card numbers (Luhn-valid 13–16 digit formats for Visa, Mastercard, AmEx, and Discover) and replaces them with tokens before any external transmission occurs.
Opens the PII Scrubber with the value from this section already filled in.
Open in the tool →PCI DSS scope and AI tools
PCI DSS Requirement 3 mandates protecting stored cardholder data. Requirement 4 requires protecting data in transit. When you send a Visa number, expiry, and CVV to an external AI service, all three requirements apply, and most AI providers are not PCI DSS certified processors. Including an uncertified AI tool in your data flow means that tool enters your PCI DSS scope, which triggers an audit obligation for the tool that the provider cannot satisfy. Consequently, the cleanest approach is to ensure cardholder data never reaches the AI provider by masking it before the request2.
This scope-extension problem is one of the most overlooked risks in payment security. When a fraud analyst pastes a transaction log containing full card numbers into ChatGPT for pattern analysis, that AI tool effectively becomes part of the cardholder data environment. The organization then needs to demonstrate that the AI provider meets all twelve PCI DSS requirements3, which is impossible with a consumer-grade AI service. Pre-scrubbing the card numbers before the paste keeps the AI tool entirely outside the cardholder data environment and eliminates this scope-creep risk.
Card formats the scrubber detects
The scrubber detects credit card numbers using the Luhn algorithm check on digit sequences of 13 to 16 digits in standard formats: spaces between groups (4242 4242 4242 4242), dashes between groups, or no separator. It catches Visa (starts with 4, 16 digits), Mastercard (51 to 55 prefix), American Express (34/37 prefix, 15 digits), Discover (6011 prefix), and JCB (35 prefix). Yet CVV codes (3 to 4 digit standalone numbers) are too short to reliably detect, so they require manual removal. Cardholder names are not detected by pattern and also require manual redaction.
The Luhn algorithm validation is what separates reliable card number detection from naive digit matching. A random 16-digit number has roughly a 1 in 10 chance of passing the Luhn check by coincidence4, so the algorithm dramatically reduces false positives compared to simple digit-counting. For payment teams that handle large transaction logs, this means the scrubber flags genuine card numbers with high confidence while ignoring order IDs, reference numbers, and other numeric fields that happen to be long digit strings.
Other PCI DSS-relevant data in AI prompts
Beyond PANs, PCI DSS-relevant data that appears in AI prompts includes payment system credentials, and API keys for payment gateways like Stripe are detected (sk_live_, sk_test_, pk_live_, pk_test_, rk_live_)5. Building on this, internal payment system hostnames (.corp, .internal) that identify payment infrastructure are detected as domain names. Furthermore, email addresses associated with cardholder accounts are caught. Combining PAN masking with credential scrubbing covers the primary PCI DSS exposure in AI prompt workflows.
Payment system logs are a particularly dense source of PCI-relevant data. A single failed transaction log entry can contain the full PAN, the cardholder email from the billing record, the Stripe API key from the payment processor config, and the internal service hostname from the infrastructure metadata. When a developer pastes this log into an AI tool to debug the failure, all four data types travel to the AI provider. The scrubber catches the PAN, email, API key, and hostname in a single pass, removing the PCI scope from the pasted content.
PCI DSS v4.0 Requirement 3 and display masking requirements
PCI DSS version 4.0, published in March 2022 and mandatory since March 20246, strengthened Requirement 3 around cardholder data display. Requirement 3.3.1 specifies that the PAN must be masked when displayed, with only the first six and last four digits visible at maximum (the format 4111 **** **** 1111). Requirement 3.3.2 prohibits displaying the full PAN except to personnel with a legitimate business need. Requirement 3.3.3 prohibits displaying the CVV, expiry date, and track data anywhere in cleartext, including logs and debugging interfaces7.
Display masking under PCI DSS v4.0 applies to cardholder data shown in any format: application UIs, log entries, support ticket fields, and administrative dashboards all fall under the requirement. When a developer pastes a dispute resolution log that shows full PAN numbers into an AI tool for analysis, that log violates Requirement 3.3.1's display restriction, and the transmission to the AI provider violates Requirement 4.2 (protecting cardholder data in transit). Masking the PAN using the scrubber before pasting satisfies the display requirement for the pasted text and eliminates the Requirement 4.2 transmission exposure simultaneously.
How PCI DSS data masking differs from PCI DSS tokenization
PCI DSS recognizes tokenization as an official scope-reduction technique under Requirement 3.5. A compliant tokenization system replaces PANs with tokens that have no exploitable relationship to the original, with the mapping stored in a certified token vault. This approach removes the tokenized card data from PCI DSS scope, because the token alone has no value to an attacker who intercepts it. The scrubber's token replacement is a data masking tool for operational sharing purposes, not a PCI-certified tokenization system. Scrubbed output is safe to share externally but does not remove the underlying PAN processing from your Cardholder Data Environment scope.
Cardholder data in payment dispute and chargeback workflows
Payment dispute workflows concentrate cardholder data at every step. A chargeback record from Visa or Mastercard's dispute portals includes the full PAN, transaction amount, merchant descriptor, dispute reason code, and the cardholder's name. Fraud analysts who use AI tools to classify dispute reasons, generate response templates, or identify dispute patterns regularly paste these records into AI tools, bringing full PAN data out of the PCI CDE and into an uncertified AI provider's environment.
Dispute management platforms such as Chargebacks911, Midigator, and Kount provide case management dashboards that aggregate chargeback records from multiple payment processors. These platforms export dispute data in CSV or JSON format for batch analysis. Before using exported dispute data in an AI tool, scrub all PAN values from the export. Dispute analysis requiring AI pattern recognition can proceed from the scrubbed version: the scrubber preserves transaction amounts, dates, dispute codes, and merchant descriptors while removing the card numbers that create PCI scope extension.
Scrubbing before sharing dispute evidence with payment processors
Merchants submitting dispute evidence to payment processor portals sometimes use AI tools to draft the narrative response and summarize supporting documents. These supporting documents (signed receipts, order confirmations, IP logs) may contain PANs, expiry dates, and cardholder contact information. Scrub PAN fields before drafting the narrative, along with any cardholder personal information in these documents. The AI receives enough context from the scrubbed documents to generate an accurate narrative, while the actual PAN never leaves your controlled document preparation environment and never reaches an uncertified external system.
Card brand tokenization programs versus local masking tools
Visa and Mastercard operate network tokenization services that replace PANs at the issuer level with network tokens for digital commerce. A network token (a 16-digit number starting with specific BIN ranges reserved for tokens) replaces the PAN for a specific merchant and device combination. The token is cryptographically bound to the merchant, the device, and the issuer, making it worthless outside that specific combination8. Network tokenization is an automatic process managed between the payment network, issuer, and the merchant's payment processor.
The scrubber's masking approach is entirely separate from card brand tokenization and does not interact with payment network systems. The scrubber operates on text that you paste (chargeback records, transaction logs, dispute documentation) and replaces PAN-formatted values with human-readable tokens like [CC_1]. This is a privacy masking tool for operational data sharing, not a payment security system. Use card brand tokenization for the payment transaction layer; use the scrubber for the operational analysis, support, and reporting layer where cardholder data appears in documents outside the payment transaction flow.
Reducing PCI DSS scope through layered controls
The most effective PCI DSS scope reduction approach combines point-to-point encryption (P2PE) for the card capture layer, network tokenization for payment processing, and display masking for the reporting and analysis layer. P2PE encrypts cardholder data from the point of interaction and ensures unencrypted PANs never reach your application server. Network tokens replace PANs in downstream transaction records. Display masking (including the scrubber for AI tool use) prevents remaining references to cardholder data in dispute and reporting workflows from traveling outside your PCI environment. These three controls together reduce the systems in scope for a PCI DSS audit to the smallest possible set.
The scrubber supplies the reporting-layer control without pulling the AI tool into your cardholder data environment. CapyToolkit runs the masking locally in your browser, so the PAN never reaches any server during scrubbing, and the variables file is the only artifact that maps tokens back to the real card numbers for your own later use.
When to use this
Use this before pasting any transaction record, payment log, card dispute document, or fraud analysis input into an AI tool when the text contains credit card numbers or payment system credentials.
Examples
Fraud analysis with transaction data
Transaction declined: card 4111111111111111 exp 12/26, billing email: [email protected], amount $4,500.
Transaction declined: card [CC_1] exp 12/26, billing email: [EMAIL_1], amount $4,500.
The transaction context (amount, date, decline reason) is preserved. PAN and email are tokenized.
Customer dispute log with multiple card formats
Dispute: card ending 4242 (full: 5500 0000 0000 0004), customer [email protected], Stripe key sk_live_xyz123
Dispute: card ending 4242 (full: [CC_1]), customer [EMAIL_1], Stripe key [STRIPE_1]
- 1.
PCI Security Standards Council, "PCI DSS v4.0 — Requirements 3 and 4," pcisecuritystandards.org, accessed June 2026. https://listings.pcisecuritystandards.org/documents/PCI-DSS-v3-2-1-to-v4-0-Summary-of-Changes-r1.pdf
- 2.
PCI Security Standards Council, "PCI DSS v4.0 Resource Hub," pcisecuritystandards.org, accessed June 2026. https://blog.pcisecuritystandards.org/pci-dss-v4-0-resource-hub
- 3.
Wikipedia, "Payment Card Industry Data Security Standard," accessed October 2026. https://en.wikipedia.org/wiki/Payment_Card_Industry_Data_Security_Standard
- 4.
Stripe, "API Key Identification," stripe.com, accessed June 2026. https://docs.stripe.com/keys
- 5.
PR Newswire, "Securing the Future of Payments: PCI SSC Publishes PCI Data Security Standard v4.0," prnewswire.com, March 31, 2022. https://www.prnewswire.com/news-releases/securing-the-future-of-payments-pci-ssc-publishes-pci-data-security-standard-v4-0--301515073.html
- 6.
WithPCI, "Requirement 3.3 — SAD Not Stored After Authorization," withpci.com, accessed June 2026. https://withpci.com/requirements/3/3.3
- 7.
The Algo, "EMV 3DS and Network Token Payment Card Tokenization," the-algo.com, accessed June 2026. https://www.the-algo.com/insights/emv-3ds-network-token-payment-card-tokenization
- 8.
PaymentBrief, "Payment Tokenization Beyond PCI Compliance," paymentbrief.com, accessed June 2026. https://paymentbrief.com/articles/payment-tokenization-beyond-pci/
CVV codes are 3–4 digit numbers, which are too short and too common to reliably detect without high false-positive rates. Remove CVV codes manually before using the scrubber.
Yes. The credit card detection uses Luhn validation to reduce false positives. Random 16-digit numbers that fail the Luhn check are not flagged as card numbers.
Stripe live keys (sk_live_, pk_live_, rk_live_) and test keys (sk_test_, pk_test_) are detected and tokenized as [STRIPE_N].
Tokens are not cardholder data under PCI DSS. If the only card-related content reaching the AI is tokens like [CC_1], that content is not in scope for PCI DSS, and the AI provider does not enter your PCI environment for that interaction.
No. Nothing is transmitted or stored. The token mapping lives only in browser memory and is cleared on tab close. CapyToolkit is not involved in any payment processing.