What to Redact on Each Type of Document

Extract text from images and redact sensitive regions locally. Tesseract WASM runs in your browser — no uploads, no server, no account needed.

What to Redact on Each Type of Document

The technique of redaction barely changes from one document to the next: cover the field on the scanned image, check that the extracted text no longer contains it, and export a flat image with no text layer underneath. What changes is the list of fields. A bank statement exposes account and routing numbers, a court filing has rules about Social Security numbers and birth dates, a medical record carries health identifiers, and a passport carries a machine-readable line that repeats most of the page.

The OCR Redactor reads the scanned page in your browser, shows the extracted text next to the image, and blanks every word your rectangle touches in both. The overview below covers the fields that recur on almost every document and the checks that apply to all of them. The entries then list what to redact on financial documents, legal filings, contracts, medical records, passports and IDs, and FOIA releases.

Before you export a redacted page

  • Every visible copy of a field account numbers and names often repeat in headers, footers and stamps
  • Text panel check a covered word shows as block characters in the extracted text
  • Handwriting and photos OCR may miss them, so cover them by their position on the image
  • Export format use the PNG or PDF export, which carries no text layer under the black boxes

Opens the Offline OCR & Document Redactor with this page's checklist shown at the top of the tool.

Open in the tool →

Fields that recur on almost every document

Most documents share a core set of identifiers, and the specific rules for each document type sit on top of that set, so a passport adds its machine-readable zone and a medical record adds health identifiers to the same base list. Starting from the shared fields means a new kind of document never begins from a blank list.

Identifiers that appear everywhere

Full names, signatures, home addresses, dates of birth, Social Security or national ID numbers, account numbers and phone numbers show up on statements, filings, forms and letters alike. Federal court filings in the United States set a useful baseline: unless the court orders otherwise, a filing may show only the last four digits of a Social Security, taxpayer or financial account number, the year of birth instead of the full date, and a minor's initials instead of the name.1

That rule applies only to court filings, yet it describes a sensible default for most other sharing too. A recipient who only needs to confirm whose account or record it is rarely needs more than the last four digits, and a year of birth confirms an age band without handing over the full date someone could use to impersonate you.

Start from what the recipient needs

Over-redacting can make a document useless, and under-redacting can breach a confidentiality duty, so decide the scope before you draw a single rectangle. Ask what the recipient must verify: a landlord checking income needs the employer and the amounts, not the account number, while a court reviewing a contract dispute may need the clause you were tempted to hide.

Writing that scope down helps more than it seems. A short list of fields to cover, made before you open the scan, keeps you consistent across a multi-page document and gives you something to check the finished pages against, instead of deciding page by page under time pressure, which is exactly when fields get missed.

Check the image and the text separately

The image and the extracted text are two outputs, and each needs its own check. The text panel shows block characters where a rectangle covered a word, which confirms the text export is clean. Handwritten notes, stamps and photos may never reach the text panel, because OCR reads them poorly or not at all, so cover those by their position on the image and confirm them by eye.

Sources
  1. 1.

    Cornell Law School Legal Information Institute, "Rule 5.2. Privacy Protection For Filings Made with the Court," law.cornell.edu, accessed October 2026. https://www.law.cornell.edu/rules/frcp/rule_5.2

Redacting Financial Documents Before Sharing

Bank statements expose account numbers, balances, and spending patterns. Sharing an unredacted bank statement for a rental application, mortgage pre-approval, or legal proceeding discloses far more financial information than the recipient requires. Most verification purposes need only specific data points such as average balance or income level, not the full transaction history and account credentials.

Financial documents also contain routing numbers, which together with account numbers are sufficient to initiate ACH transfers in many jurisdictions. Redacting routing numbers, full account numbers, and SSNs before sharing any financial document limits exposure to credential fraud, regardless of the recipient's trustworthiness. The OCR Redactor processes all redactions locally with no upload.

What this page covers

  • Account number leave only the last 4 digits, per the worked example
  • Routing number named in the worked example
  • Transaction descriptions named in the worked example
  • SSN and home address named in the worked example for a loan application statement

Opens the Offline OCR & Document Redactor with this section's checklist shown at the top of the tool.

Open in the tool →

What to redact on bank statements and tax returns

Bank statements typically display the full account number on every page header. Redact all but the last four digits, the widely accepted convention for account verification. Routing numbers appear alongside account numbers on check images within statements in both fractional form and the nine-digit MICR form along the bottom of the check.1 The routing number alone is semi-public information because it identifies the bank rather than the specific account, but combining it with even a partial account number increases the risk of unauthorized ACH initiation, so covering both is the safer practice.2

Tax return and income document redaction priorities

Transaction descriptions can reveal sensitive spending patterns, ongoing subscriptions, or relationships you prefer not to disclose.3 For income verification, cover all transactions and leave only the opening balance, closing balance, and summary figures. Tax returns contain SSNs or EINs, full name, address, and complete income breakdown.4 Consequently, redact SSN, EIN, address, and any income sources not required for the stated purpose before sharing any tax document. CapyToolkit processes each page locally, so sensitive financial data never leaves your device during this workflow. For tax returns used in a divorce or custody proceeding, the required disclosure scope is typically limited to income and deduction totals, not the full attachment schedules that reveal business relationships and investment holdings.

Pay stubs and financial records for employment verification

Pay stubs contain employer name, employee address, Social Security number, gross income, net income, year-to-date totals, and payroll deduction details.5 For a rental application that requires income verification, the employer name, gross income, and year-to-date figures are typically the only required fields. Redact the SSN, home address, and any deduction detail that reveals medical, retirement, or other benefit elections not relevant to the inquiry. Building on this, if the purpose is only to confirm employment, further redact the income amounts entirely and leave only the employer name and employment dates visible. Adjust the disclosure to exactly what the recipient has a legitimate need to see.

Workflow for financial document redaction

Scan or photograph the financial document and save as JPEG. Drop it onto the OCR Redactor drop zone. After OCR completes, identify every sensitive field: account numbers, routing numbers, SSNs, EINs, balances (if not required), and address fields. Draw rectangles over each. Confirm in the text panel that those strings no longer appear as readable text in the export. Export the PNG. For multi-page statements, process each page individually and combine the exported PNGs into a single PDF using a local PDF tool. Yet the most important verification step remains visual inspection of the final PNG before sending, not just the text panel review, because OCR may not detect every field accurately.

Redacting brokerage and investment account statements for legal proceedings

Brokerage statements expose more than a simple bank balance. A single monthly statement typically reveals account number, account type, total portfolio value, individual holding names and quantities, cost basis per position, realized gains and losses, dividend and interest income, and transaction history for the period. In a divorce proceeding, personal injury suit, or business dispute where asset disclosure is required, the minimum necessary scope for a specific proceeding may require only the total portfolio value and certain income figures, not the full holding breakdown that reveals investment strategy and unrealized positions.

Redact individual holding names and quantities for positions not relevant to the proceeding, leaving only the portfolio total and required income lines. For tax-related proceedings requiring capital gains verification, leave realized gains and losses visible but cover unrealized position details and cost basis for non-relevant holdings. Draw rectangles over each individual position row in the holding detail section and each transaction line outside the relevant period, leaving summary totals visible. The OCR Redactor supports narrow rectangles that cover specific rows without affecting adjacent summary rows above and below.

Financial document redaction for due diligence disclosures

Business due diligence processes require disclosing financial records to prospective buyers, investors, or lenders who have signed an NDA but who are not yet authorized to see all details. A prospective acquirer performing due diligence on a target company may need to verify revenue and EBITDA figures without receiving the full customer list embedded in the revenue breakdown. For these situations, redact individual customer names and transaction amounts in accounts receivable aging reports while leaving the aggregate total visible. Redact individual vendor names and payment amounts in accounts payable records while leaving total liabilities visible. This selective disclosure allows financial verification without premature disclosure of customer relationships and vendor agreements that remain commercially sensitive until the transaction closes.

Checking for financial identifier patterns in the OCR text panel after redaction

After placing all redaction rectangles, use the OCR text panel to verify that no financial identifier patterns remain visible in the export. Search the text panel for nine-digit numbers (SSN format: XXX-XX-XXXX), sequences of eight or more consecutive digits (account numbers), and the routing number format (nine digits at the start of a line, which is the standard position on bank letterhead). The OCR text extraction may not capture every number accurately, particularly on older bank statement designs with compressed fonts or light toner, so visual inspection of the image remains the primary verification step. The text panel search supplements visual review by catching identifiers in regions where your visual sweep may have moved quickly.

For brokerage statements with CUSIP numbers (nine-character alphanumeric identifiers for each security holding), search the text panel for nine-character strings beginning with a digit followed by two letters. CUSIP numbers identify specific securities and may require redaction when the holding itself is sensitive. For statements with ISIN codes (12-character international security identifiers beginning with a two-letter country code), search for 12-character uppercase strings beginning with two letters. Both identifier types appear in holding detail tables and are small enough that OCR sometimes merges them with adjacent text, making visual confirmation of their absence in the exported image the necessary final check.

Organizing multi-document financial redaction sessions

Financial disclosures for proceedings and due diligence typically involve multiple document types processed in a single session: bank statements, brokerage statements, tax returns, pay stubs, and loan documents. Before starting, list every document and the specific fields to redact in each, based on the disclosure scope defined by the proceeding rules or the NDA. Process one document type at a time, applying the same field checklist to each instance of that type before moving to the next type. This batch-by-type approach reduces the risk of applying the wrong redaction scope to a document, because all instances of a given document type share the same redaction rules and your attention stays calibrated to that type's specific field layout during the processing run.

For very large productions, group documents by type before you open the tool at all. CapyToolkit's OCR Redactor redacts each image you drop in, so a written field checklist per document type keeps the disclosure scope consistent across every instance and makes it easy to spot a page that drifted outside its rules. Reconfirm the NDA or proceeding scope once at the start of the session, then apply the same rectangles to each matching page without re-deciding field by field, beginning with mask account numbers past four digits on every statement.

When to use this

Use this tool when submitting financial documents for rental applications, mortgage approvals, legal proceedings, or business due diligence where you want to limit disclosure to exactly the information the recipient requires.

Examples

Submitting a bank statement for a rental application

Redact account number (leave last 4 digits), routing number, and all individual transaction descriptions. Leave opening and closing balances and the account holder name visible. Export the PNG.

Providing a tax return for a mortgage application

Cover SSN, home address, and all income lines not required by the lender. Leave only the adjusted gross income total and the year visible. Export the PNG and share it in place of the full return.

Sources
  1. 1.

    Cornell Law School LII, "12 CFR Appendix A to Part 229 — Routing Number Guide," law.cornell.edu, accessed June 2026. https://www.law.cornell.edu/cfr/text/12/appendix-A_to_part_229

  2. 2.

    NACHA, "ACH File Details," achdevguide.nacha.org, accessed June 2026. https://achdevguide.nacha.org/ach-file-details

  3. 3.

    Consumer Financial Protection Bureau, "Personal Financial Data Rights Reconsideration ANPR," consumerfinance.gov, August 2025. https://files.consumerfinance.gov/f/documents/cfpb_personal-financial-data-rights-reconsideration_anpr_2025-08.pdf

  4. 4.

    Internal Revenue Service, "1040 Instructions (2025)," irs.gov, accessed June 2026. https://www.irs.gov/instructions/i1040gi

  5. 5.

    U.S. Department of Labor, "Fact Sheet #21: Recordkeeping Requirements under the FLSA," dol.gov, July 2008. https://www.dol.gov/agencies/whd/fact-sheets/21-flsa-recordkeeping

FAQ

Yes. Draw individual rectangles over each transaction line you want to redact. The tool supports multiple non-overlapping rectangles on the same image. Each rectangle independently removes the covered words from the text panel. Take your time placing rectangles precisely over each transaction description.

Yes. PCI DSS requires masking all but the last four digits of payment card numbers in displays and on receipts. Most regulated financial disclosure contexts follow the same convention for bank accounts. Showing the last four digits confirms the account without exposing the full credential.

Routing numbers are semi-public; they identify the bank rather than the account and are published by the Federal Reserve. The account number combined with the routing number is the actual credential for ACH initiation. Redacting the account number provides the meaningful protection; routing number redaction is an additional precaution.

Yes. Some lenders require unredacted statements for underwriting purposes. In that case, submit the unredacted version only to the lender through a secure channel, and use redacted versions for any other parties involved in the application process such as brokers, co-signers, or property agents.

Process the pages that contain sensitive fields first. Identify which pages have account numbers, routing numbers, and SSNs (usually the header page and any check image pages). Pages with only transaction histories can be processed more quickly by drawing a single large rectangle over the transactions section if you want to redact all of them at once. CapyToolkit handles each page individually, so you control exactly which fields are visible in the final export.

Redacting NDA and Contract Terms Before Distribution

NDA redaction protects confidential terms before distribution. Sharing a contract with a third party for reference, dispute resolution, or regulatory disclosure often requires redacting clauses that the counterparty marked as confidential, that reveal pricing or intellectual property terms not relevant to the third party's inquiry, or that identify other parties bound by separate non-disclosure obligations.

The challenge with contract redaction is defining scope correctly. Over-redacting may make a document useless for its intended purpose; under-redacting may breach a confidentiality obligation to another party. The OCR Redactor provides the technical means for precise field-level redaction on scanned contracts with no upload, making it suitable for privileged document workflows in legal and compliance contexts.

What this page covers

  • Confidential information definition named in the worked example
  • Compensation terms named in the worked example
  • Royalty rate and license fee schedule named in the second worked example
  • Minimum commitment thresholds named in the second worked example

Opens the Offline OCR & Document Redactor with this section's checklist shown at the top of the tool.

Open in the tool →

What contract clauses typically require redaction

Standard NDA redaction targets three categories of information.1 First, the confidential information definition clause, which defines the scope of what is protected, may itself be marked as confidential in some NDAs. Second, compensation, pricing, and payment terms reveal financial relationships that parties routinely redact in third-party disclosures. Third, term and renewal dates may reveal the duration of a business relationship that the parties intended to keep private.

Party identification and third-party confidentiality

Party identification beyond what the receiving party already knows may require redaction when the contract identifies sub-contractors, licensors, or investors by name whose involvement the original parties agreed to keep confidential. Consequently, review the contract's own confidentiality and non-disclosure provisions before deciding which sections to cover. CapyToolkit's local processing ensures that privileged contract content stays on your device throughout the redaction workflow, which matters when the contract itself contains the sensitive terms you are protecting.

Technical approach to contract redaction

Scanned contracts often span multiple pages with dense text layouts. Scan or export each page at 300 DPI or better, the resolution Tesseract documentation names as its working baseline2, and drop each page image separately into the OCR Redactor. After OCR extracts the text, draw rectangles precisely over the clause text, heading, and any marginal annotations that form part of the confidential section. The text panel updates to show block characters replacing the covered text, providing a preview of the de-identified text export. Building on this, drawing a rectangle over a section heading alone does not redact the section body; draw rectangles over both the heading and the full body text of any clause you need to cover. Review the text panel carefully for any clause text that may have wrapped into adjacent paragraphs.

Exporting and using the redacted contract

Export the redacted page as a PNG or PDF. For multi-page contracts, export each page, then combine the pages into a single PDF using any local PDF tool. Label the combined document as a redacted version with the redaction basis noted in cover correspondence (such as "Redacted per Section 12 confidentiality obligations"). Retain the original unredacted contract securely. In litigation or regulatory inquiries, maintain a redaction log identifying each covered section by clause number and the basis for the redaction, as privilege logs and redaction indices are standard practice in civil discovery and regulatory productions.3 A redaction log records each redaction made and its basis, providing a defensible record if the redaction decisions are later challenged.4 Yet the log is your responsibility; the tool provides the image export only.

Identifying third-party confidentiality obligations within an NDA before redacting

Many NDAs contain provisions protecting information belonging to third parties, not just the two signing parties. Vendor agreements, licensing terms, and technical specifications shared under a prior NDA often appear embedded in a later contract without explicit marking. Before redacting for disclosure to a new party, review the agreement for any clause that references obligations to a third party, incorporates terms from another agreement by reference, or uses language such as "confidential information received from" followed by a third-party name. Each such clause may require redaction even if the original signing parties would otherwise consent to disclosure.

The confidential information definition in most NDAs is the controlling clause for redaction scope.5 Broad definitions (any information disclosed in connection with the business relationship) require conservative redaction of nearly all substantive terms. Narrow definitions (only information marked CONFIDENTIAL in writing) allow disclosure of unmarked terms without redaction. Read the definition clause first, then review the document with that scope in mind before drawing any rectangles. Redacting based on a misreading of the definition clause exposes the redacting party to liability if the disclosure violates the actual contractual obligation.

Handling counterparty identity redaction in multi-party agreements

Contracts involving more than two parties require careful identity redaction when one party receives a copy that should not reveal the identities of other parties. Joint venture agreements, consortium arrangements, and multi-party licensing deals frequently include schedules listing all participating entities by name and address. For disclosures to a party who is aware of their own participation but should not know the identities of co-participants, redact all party names and addresses other than the recipient's own. Leave signature blocks, notice addresses, and obligation clauses intact to allow the recipient to understand their own rights and obligations without revealing the full participant list.

Redacting contract schedules, exhibits, and incorporated documents

Contract schedules and exhibits often contain the most sensitive substantive content: pricing schedules, technical specifications, IP disclosures, and personnel lists appear as attachments rather than body text precisely because they are detailed and sensitive. Redacting the main body of an NDA without addressing its schedules leaves the most sensitive content exposed. For each exhibit or schedule referenced in the main agreement, assess whether the exhibit's contents fall within the redaction scope defined by the agreement's own confidentiality provisions before deciding what to cover.

Incorporated documents create a related challenge. When an NDA states "the parties' obligations under the Master Services Agreement dated [date] are incorporated herein," the MSA's terms become part of the NDA for redaction purposes. If the MSA contains pricing or IP terms that require redaction, those terms require redaction when the NDA is disclosed even though they appear in a separate document. Maintain a list of all incorporated documents referenced in any contract you process, and assess each for redaction scope alongside the primary agreement.

Building a consistent redaction checklist for recurring contract types

Recurring document types, such as standard vendor NDAs, employment agreements, or SaaS subscription terms, contain predictable field structures that appear in the same positions across different counterparty versions. For each recurring contract type, develop a checklist of fields that always require redaction: the counterparty's legal name and address, any compensation or fee terms, the confidential information definition clause text, expiry or term dates if competitively sensitive, and any field specifically marked as confidential within the document. Apply the same checklist to every instance of that contract type, then supplement with document-specific items discovered during review. This systematic approach reduces the risk of missing sensitive fields that appear at consistent positions and eliminates the need to re-assess scope from scratch on each instance.

Keeping the checklist in a shared location helps teams apply the same scope across reviewers. CapyToolkit's OCR Redactor redacts exactly the regions you mark, so the checklist is what guarantees consistency from one contract to the next rather than the tool guessing at confidential terms. Update the list whenever a new clause type appears in a revised template, because a stale checklist is the most common reason a sensitive field slips through on a high-volume redaction run, so hide confidential contract clauses by hand using the same field list every time.

When to use this

Use this tool when producing contracts in litigation or regulatory proceedings, when sharing agreements with auditors or investors who are not party to all confidentiality provisions, or when providing contract extracts for reference in negotiations where only certain terms are relevant.

Examples

Sharing an NDA with an auditor who needs to verify the agreement exists

Redact the confidential information definition, compensation terms, and names of any non-audit parties. Leave the execution date, parties' identities relevant to the audit, and term length visible. Export the PNG.

Producing a licensing agreement in litigation with pricing terms asserted as trade secret

Draw rectangles over the royalty rate, license fee schedule, and any minimum commitment thresholds. Log each redacted section on the privilege log with the trade secret basis. Export each page as PNG and combine into a production PDF.

Sources
  1. 1.

    Association of Corporate Counsel, "The Key Elements of a Great NDA," acc.com, 2020. https://www.acc.com/sites/default/files/program-materials/upload/DLA%20-%20The%20Key%20Elements%20of%20a%20Great%20NDA.pdf

  2. 2.

    Tesseract OCR Documentation, "Improving the Quality of the Output," tesseract-ocr.github.io, accessed June 2026. https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html

  3. 3.

    American Bar Association, "Four Types of Privilege Logs That Litigators Need to Know About," americanbar.org, 2024. https://www.americanbar.org/groups/law_practice/resources/law-technology-today/2024/four-types-of-privilege-logs-that-litigators-need-to-know-about/

  4. 4.

    Bill Gallivan, "Why Redaction Logs Matter in eDiscovery and Document Disclosure," digitalwarroom.com, 2024. https://www.digitalwarroom.com/blog/why-redaction-logs-matter

  5. 5.

    Sergei Tokmakov, "Definition of Confidential Information," terms.law, accessed June 2026. https://terms.law/NDA/clause-library/definition-of-confidential-information/

FAQ

Yes. If opposing counsel challenges a redaction in discovery, the court may order in camera review of the unredacted version and rule on whether the redaction is valid. Courts regularly order production of unredacted documents when they find the redaction basis insufficient. Always retain the original.

Regulatory agencies such as the SEC, FDA, and NLRB each have their own rules on contract redactions in filings. Some allow redaction of trade secrets and confidential commercial information; others require full disclosure with a specific exemption request. Check the applicable agency's rules before filing a redacted version.

Draw a small rectangle precisely over the word or phrase within the paragraph. The tool supports rectangles of any size. The rectangle can be narrow enough to cover a single dollar amount or a specific name within a sentence. Zoom in on the image if the contract text is small to place the rectangle accurately.

The image redaction is independent of OCR quality. The rectangle covers the image pixels regardless of what OCR extracted from that region. A garbled text export does not affect the visual redaction. For the text export, you can manually correct garbled text in the text panel before downloading it.

No. CapyToolkit generates the exported PNG from a browser canvas element, which contains no metadata inherited from the source scan. The original JPEG or PNG you loaded may contain EXIF data, but the exported redacted PNG does not carry it. No document metadata, creation dates, or author information appears in the browser canvas export.

Redacting Medical Records: Removing PHI from Scanned Documents

Medical records contain the most sensitive personal data possible. Improper redaction of scanned health records, specifically adding a black rectangle on top of a PDF while leaving the underlying text layer intact, is a well-documented failure, because the original text stays in the file and can be uncovered by deleting the overlay1. The OCR Redactor exports raster images with no underlying text layer, eliminating this specific risk category from any document it processes.

Processing happens entirely in the browser. No file leaves your device at any point during OCR or redaction. For healthcare professionals, compliance staff, and patients requesting copies of their records before sharing them with third parties, this zero-upload approach removes the need for a business associate agreement with any tool vendor.

What this page covers

  • All 18 HIPAA Safe Harbor identifiers must be visible and redacted in the scanned record before it is produced
  • Text layer the exported PNG has none, so identifiers can't be recovered programmatically

Opens the Offline OCR & Document Redactor with this section's checklist shown at the top of the tool.

Open in the tool →

What PHI requires redaction in scanned records

HIPAA's Safe Harbor de-identification method defines 18 protected health information identifiers that must be removed before a health document can be shared without patient authorization2. These include name, address, all date elements more specific than year except for ages over 89, phone numbers, fax numbers, email addresses, Social Security numbers, medical record numbers, health plan beneficiary numbers, account numbers, certificate or license numbers, vehicle identifiers, device identifiers, URLs, IP addresses, biometric identifiers, full-face photographs, and any other unique identifying number or code.

Systematic page-by-page identifier review

Scanned records often contain all 18 identifier categories across multiple pages. Consequently, drawing redaction rectangles over each identifier in the image removes the corresponding words from the exported text file simultaneously. CapyToolkit's OCR Redactor processes each page individually, so you can methodically work through a multi-page record one sheet at a time, confirming every identifier category is covered before moving to the next page.

Why standard PDF redaction fails for medical records

Most PDF editors that offer a redaction feature add an opaque rectangle to the visual layer of the document. Yet the text layer remains intact in the PDF structure1. Opening the redacted PDF in a text editor or using a PDF parser extracts the covered text in seconds. Several publicized HIPAA breaches resulted directly from this type of improper redaction. Building on this documented failure pattern, the safe approach for any health record shared outside a covered entity is a flat raster image that contains only pixels, no text layer, no metadata carrying patient identifiers, and no structure that any parser can traverse to recover covered content.

Workflow for medical record redaction

Scan or photograph the medical record and export it as a JPEG or PNG. Drop it onto the OCR Redactor drop zone. After OCR completes, draw rectangles over each of the 18 PHI identifiers visible in the scan. Pay particular attention to small-font identifiers in page headers, footers, and patient ID boxes that OCR may have partially missed. Visually inspect the redacted image before exporting to confirm no identifiers remain visible. Export using Download PNG, which produces a clean raster image. Yet the tool is a technical aid; organizational HIPAA compliance requires documented policies, staff training, and audit trails beyond what any single tool provides. Document each redaction decision for your compliance records.

HIPAA's two de-identification methods and their document requirements

HIPAA's Privacy Rule at 45 CFR §164.514(b) provides two recognized paths to de-identification3. Safe Harbor requires removing all 18 enumerated identifiers and any other information the covered entity has reason to believe could identify the individual. Expert Determination allows a qualified statistician to certify that re-identification risk is very small, without necessarily removing all 18 identifiers, provided the analysis is documented and retained.

Safe Harbor identifier categories for document redaction

The 18 Safe Harbor identifiers include names, geographic subdivisions smaller than a state, dates other than year for individuals over 89, telephone and fax numbers, email addresses, Social Security numbers, medical record numbers, health plan beneficiary numbers, account numbers, certificate or license numbers, vehicle identification numbers, device serial numbers, URLs, IP addresses, biometric identifiers, full-face photographs, and any other unique identifying code. When redacting a medical record scan, walk through each category systematically before export. A single missed identifier such as a fax number in a page header or a beneficiary number in a routing field means the document does not meet Safe Harbor standard.

Multi-page medical record redaction and final document assembly

Medical records frequently span multiple pages: a face sheet, physician notes, lab results, and a discharge summary may each be a separate scanned page. Export each page as an individual PNG from the OCR Redactor after completing all redactions, using a consistent naming convention such as patient_record_p01.png and patient_record_p02.png so pages assemble in order.

Assembling redacted pages into a single PDF

After exporting all pages as PNGs, bundle them into a single PDF using a local tool. ImageMagick's convert command accepts a glob of PNG filenames4 and produces a multi-page raster PDF in filename sort order. Alternatively, macOS Preview accepts multiple PNGs dragged into the thumbnail sidebar and exports a PDF via File > Export as PDF. The resulting PDF contains only raster image data with no searchable text layer, unlike OCR-generated searchable PDF output, which keeps the page image alongside a hidden searchable text layer5. This preserves the redaction guarantee established by each individual OCR Redactor export.

Assembling locally preserves the privacy benefit of the per-page redaction. CapyToolkit's OCR Redactor keeps each page a flat raster image, so the combined PDF carries no recoverable text layer even after you merge the pages into one file, the approach behind redacting PHI one scanned page at a time. Choose whichever bundling tool you already trust on your machine, because the merge step introduces no new upload and never sends the pages to a remote server. The redaction guarantee therefore holds through to the final shared document.

When to use this

Use this tool when disclosing health records in response to legal requests, when sharing records for research under a limited data set authorization, or when preparing de-identified records under the Safe Harbor method.

Examples

Responding to a subpoena for patient records

Redact all 18 PHI identifiers visible in the scanned records before producing them. Export each page as a PNG and log the redacted fields. Retain the originals under seal as required by the court order.

Preparing a de-identified dataset for a quality improvement study

Process each patient record image, redact all PHI, and export PNGs. The exported images contain no text layer from which identifiers could be recovered programmatically, satisfying the Safe Harbor technical requirement.

Sources
  1. 1.

    Wikipedia, "Redaction," accessed October 2026. https://en.wikipedia.org/wiki/Redaction

  2. 2.

    HHS, "Guidance Regarding Methods for De-identification of Protected Health Information," hhs.gov, accessed June 2026. https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html

  3. 3.

    Cornell Law School Legal Information Institute, "45 CFR § 164.514 - Other requirements relating to uses and disclosures of protected health information," law.cornell.edu, accessed October 2026. https://www.law.cornell.edu/cfr/text/45/164.514

  4. 4.

    ImageMagick, "Command-line Processing," imagemagick.org, accessed June 2026. https://imagemagick.org/command-line-processing/

  5. 5.

    Tesseract OCR Documentation, "Frequently Asked Questions," tesseract-ocr.github.io, accessed June 2026. https://tesseract-ocr.github.io/tessdoc/FAQ.html

FAQ

The tool provides technically sound redaction (raster export, no text layer, no upload). HIPAA compliance also requires organizational safeguards: access controls, workforce training, audit trails, and breach notification procedures. The tool addresses the technical redaction step but does not replace the organizational program.

Yes. OCR accuracy depends on scan quality. Handwritten notes, stamps, and very small fonts can produce garbled or missing text in the extraction panel. Always review the image visually and draw rectangles over all visible PHI, not just the fields that appear in the text panel. The redaction layer operates on the image, not only on OCR-detected text.

HIPAA does not specify separate retention rules for redacted versions. Your organization's records retention policy and applicable state law govern how long any version of a medical record must be kept. The OCR Redactor has no retention function; it produces an exported file that you manage according to your own policies.

Currently the CapyToolkit OCR Redactor processes one image at a time. For multi-page records, process each page separately and export each redacted page as a PNG. Combining the exported PNGs into a PDF requires a separate step using any PDF assembly tool, all of which can be run locally without uploads.

The exported PNG and PDF contain no EXIF or document metadata because they are generated from a clean canvas in the browser, not from the original file. The original image file remains unchanged on your device. The exported redacted copy carries no identifying metadata from the source.

HIPAA-Compliant Document Redaction: PHI Removal Without the Cloud

HIPAA violations most often involve improper disclosure and inadequate technical safeguards. The Office for Civil Rights enforcement record treats improper disclosure of protected health information as a central technical safeguard failure covered by resolution agreements and civil monetary penalties.1 Adding a black rectangle on a PDF while leaving the underlying searchable text layer intact does not satisfy HIPAA's requirement to de-identify or redact protected health information before disclosure.

The OCR Redactor exports flat raster images with no searchable text layer. Processing occurs entirely in the browser using WebAssembly, which means no patient data transmits to any server at any point. Because no electronic PHI leaves the local device, no business associate agreement with the tool provider is required under HIPAA's business associate rules for tools that process data locally without storing it.

What this page covers

  • All 18 Safe Harbor identifiers the first condition for Safe Harbor de-identification
  • Limited data set exceptions dates, county or zip-level geography, and ages may remain visible depending on the data use agreement

Opens the Offline OCR & Document Redactor with this section's checklist shown at the top of the tool.

Open in the tool →

HIPAA's 18 PHI identifiers in scanned documents

HIPAA's Safe Harbor de-identification standard requires removal of 18 categories of identifiers from health information.2 Scanned documents typically contain all of them across insurance forms, lab reports, referral letters, and admission paperwork. Name and address fields appear on every page header. Dates more specific than year appear in treatment timelines, lab dates, and insurance billing records. Full-face photographs embedded in patient ID cards and biometric identifiers captured by signature pads also count as PHI categories under the current Safe Harbor enumeration.

Locating identifiers across multi-page health records

Medical record numbers, health plan beneficiary IDs, and Social Security numbers appear on insurance forms. Drawing redaction rectangles over each visible identifier on the image removes the corresponding words from the exported text file via bounding box overlap detection, providing both visual and text-layer redaction in one step. Because CapyToolkit processes each page individually, you can work through a multi-page record systematically, covering every identifier category before moving to the next page. A single missed identifier such as a fax number in a page header or a beneficiary number in a routing field means the document does not meet Safe Harbor standard, so methodical page-by-page review is essential.

Why PDF text-layer redaction fails HIPAA requirements

Several HIPAA enforcement actions cited improper PDF redaction specifically. In those cases, covered entities produced documents to plaintiffs' counsel or government investigators with black boxes drawn over text in PDF viewers, but the underlying text remained recoverable by selecting all content or parsing the PDF file structure. OCR-produced text layers embedded in scanned PDFs are equally vulnerable. Building on this documented failure pattern, the safe approach for any health record shared outside a covered entity is a flat raster image that contains only pixels, no text layer, no metadata carrying patient identifiers, and no structure that any parser can traverse to recover covered content.

Workflow for HIPAA-context redaction

Scan the health document at 300 DPI, the industry-standard resolution for document scanning, and save the JPEG locally.3 Drop it onto the OCR Redactor drop zone. After OCR completes, draw rectangles over each of the 18 PHI identifiers visible in the scan. Pay particular attention to small-font identifiers in page headers, footers, and patient ID boxes that OCR may have partially missed. Visually inspect the redacted image before exporting to confirm no identifiers remain visible. Export using Download PNG, which produces a clean raster image. Yet the tool is a technical aid; organizational HIPAA compliance requires documented policies, staff training, and audit trails beyond what any single tool provides. Document each redaction decision for your compliance records.

HIPAA enforcement actions involving improper document redaction

The Office for Civil Rights has cited improper disclosure and inadequate technical safeguards in multiple HIPAA enforcement settlements. In several cases, covered entities used PDF markup tools to draw black boxes over PHI in scanned records, then produced the PDFs to plaintiffs or government investigators without verifying that the underlying text remained inaccessible. Recipients extracted the covered text using standard PDF reader copy functions or command-line PDF parsing tools and demonstrated that the redaction had failed. Settlements in these cases included civil monetary penalties and corrective action plans requiring staff training on proper redaction techniques.1

The technical lesson from these enforcement actions is clear: any redaction method that leaves a text layer in the output PDF file does not satisfy HIPAA's requirement to de-identify or redact PHI before disclosure. OCR-based raster redaction eliminates the text layer by design, because the output is a flat PNG or image-only PDF generated from a browser canvas with no text data structure. Sharing a raster-exported redacted record eliminates the specific failure mode present in all cited enforcement cases.

Redaction for the Expert Determination method alongside Safe Harbor

HIPAA provides two de-identification methods: Safe Harbor (removing all 18 listed identifiers) and Expert Determination (a qualified statistician certifies that re-identification risk is very small).2 Most organizations use Safe Harbor because it does not require retaining a statistician. For records that cannot meet Safe Harbor requirements due to necessary retention of some quasi-identifiers (zip codes, dates needed for clinical context, ages over 89), Expert Determination provides an alternative path. In that case, redaction applies only to identifiers the expert determines create re-identification risk given the intended disclosure context, which may be fewer or different identifiers than the Safe Harbor list. Document which method applies to each redaction batch and retain the expert's certification or Safe Harbor determination in your compliance records.

Documenting the redaction process for HIPAA audit purposes

HIPAA's Security Rule requires covered entities to maintain documentation of security safeguard decisions for six years from the date of creation or the date it was last in effect.4 For document redaction workflows, this means retaining records of what was redacted, who performed the redaction, when it occurred, and what rule or standard justified the redaction decision. At minimum, maintain a log containing the original document identifier, the date of redaction, the name or role of the person performing the redaction, the disclosure recipient, the basis for disclosure (such as a court order, patient authorization, or research data use agreement), and the fields redacted.

For organizations with high-volume redaction workflows, a simple spreadsheet maintained in a HIPAA-compliant storage location satisfies this documentation requirement without specialized software. Each redaction session produces one row. The documentation serves both audit and incident response purposes: if a redaction error is discovered after disclosure, the log identifies which staff member performed the redaction and which document was affected, enabling targeted corrective action rather than a broad program review.

Minimum necessary standard applied to health record redaction scope

HIPAA's minimum necessary standard applies to each specific disclosure purpose.5 For a disclosure to a patient's attorney in personal injury litigation, the minimum necessary scope may include all clinical records relevant to the claimed injury. For a disclosure to a quality improvement team analyzing aggregate outcomes, the minimum necessary scope may exclude all 18 PHI identifiers. For an insurance company verification request, the minimum necessary scope is typically limited to the specific dates of service and diagnosis codes the insurer requested. Before drawing any rectangle, identify the disclosure purpose and the minimum necessary scope for that purpose. Redact everything outside that scope, not just the identifiers that seem most obviously sensitive.

Applying the standard at the rectangle level keeps the disclosure within bounds. CapyToolkit's OCR Redactor redacts exactly the regions you draw over and does not auto-classify what is sensitive, so the minimum necessary scope is set entirely by your review of the disclosure purpose. Before exporting, recheck the image against that purpose and cover any field outside the authorized scope, because an over-broad disclosure is itself a HIPAA violation regardless of how carefully the obvious identifiers were treated, which is why HIPAA redaction without a BAA stays entirely on your device.

When to use this

Use this tool when disclosing health records in response to legal requests, when sharing records for research under a limited data set authorization, or when preparing de-identified records under the Safe Harbor method.

Examples

Responding to a subpoena for patient records

Redact all 18 PHI identifiers visible in the scanned records before producing them. Export each page as a PNG and log the redacted fields. Retain the originals under seal as required by the court order.

Preparing a de-identified dataset for a quality improvement study

Process each patient record image, redact all PHI, and export PNGs. The exported images contain no text layer from which identifiers could be recovered programmatically, satisfying the Safe Harbor technical requirement.

Sources
  1. 1.

    HHS Office for Civil Rights, "Resolution Agreements," hhs.gov, accessed June 2026. https://www.hhs.gov/hipaa/for-professionals/compliance-enforcement/agreements/index.html

  2. 2.

    HHS, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule," hhs.gov, accessed June 2026. https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html

  3. 3.

    "Scanning resolution: the magic 300 dpi," wisetrend.com, accessed June 2026. https://www.wisetrend.com/scanning-resolution-the-magic-300-dpi/

  4. 4.

    HHS, "45 CFR 164.316 — Policies and procedures and documentation requirements," ecfr.gov, accessed June 2026. https://www.ecfr.gov/current/title-45/part-164/subpart-C/section-164.316

  5. 5.

    HHS, "45 CFR 164.502(b), 164.514(d) — Minimum Necessary Standard," ecfr.gov, accessed June 2026. https://www.ecfr.gov/current/title-45/part-164/subpart-C/section-164.502

FAQ

The tool provides technically sound redaction (raster export, no text layer, no upload). HIPAA compliance also requires organizational safeguards: access controls, workforce training, audit trails, and breach notification procedures. The tool addresses the technical redaction step but does not replace the organizational program.

No. The OCR Redactor processes data locally in your browser using WebAssembly. No electronic PHI is transmitted to or stored by any external service at any point. Because no PHI leaves your device, the tool does not function as a business associate under HIPAA's definitions.

Safe Harbor requires removal of all 18 listed identifiers plus certification that the covered entity has no actual knowledge that the remaining information could identify the individual. Redacting the 18 fields satisfies the first condition. The certification requires a covered entity official with knowledge of the specific dataset to make the determination.

A limited data set may retain some dates, geographic data at the county or zip code level, and ages. Redact only the identifiers not permitted in a limited data set as specified in the data use agreement. The tool redacts exactly the fields you draw over; it does not auto-classify what to redact.

No. The image redaction layer operates independently of OCR results. Drawing a rectangle on the image covers that region visually regardless of whether OCR detected text there. OCR results affect only the text panel export. Always visually inspect the image to ensure all identifiers are covered, not just those appearing in the text panel.

Redacting Passports and Government IDs Before Sharing

Passport scans are a primary target for identity theft. Sharing an unredacted passport scan over email, messaging apps, or cloud storage exposes the document number, nationality, full name, date of birth, and machine-readable zone: all the information needed to impersonate someone in certain identity verification contexts.

Redacting a government ID scan before sharing requires covering at minimum the document number, the photograph, and any biometric data. What you leave visible depends on the recipient's legitimate need. A landlord verifying age needs only the date of birth; they do not need the document number or photo. Drawing rectangles over unnecessary fields and exporting the raster PNG provides a legally adequate copy for most verification purposes while limiting exposure.1

What this page covers

  • Document number named in the worked example as a field to cover
  • Photograph named in the worked example as a field to cover
  • MRZ block the two-line machine-readable zone at the bottom of the photo page
  • License number and home address the driver's license equivalent, redacted while name, date of birth, and photo stay visible

Opens the Offline OCR & Document Redactor with this section's checklist shown at the top of the tool.

Open in the tool →

What to redact on a passport or government ID

Passport document numbers follow a specific pattern: in the US Next Generation Passport book, the passport number begins with a letter followed by eight numbers, which makes it easy to pick out of a scan and worth covering everywhere it appears.2 The machine-readable zone at the bottom of the passport photo page encodes name, nationality, document number, date of birth, and expiry in a scannable format that software can extract in milliseconds. Cover the MRZ entirely when sharing a passport scan unless the recipient specifically requires it for verification.

Photograph, MRZ, and document number coverage

Dates of birth and expiry dates reveal age and validity period respectively. The photograph enables facial recognition correlation. Depending on the purpose of sharing, you may need to redact the photo, the MRZ, the document number, or all three. CapyToolkit's OCR Redactor covers each region with opaque rectangles in the exported raster image, so no software can recover the original content from the shared file.

Driver's license and national ID redaction

Driver's licenses contain a different but equally sensitive set of identifiers: license number, address, date of birth, and in some jurisdictions a digital fingerprint indicator or organ donor code. Sharing a license scan for age verification purposes requires only the date of birth to be visible. Cover the license number, home address, and any biometric indicator for a minimal disclosure that serves the age verification purpose without creating an identity theft risk. Building on this, national ID cards and residence permits carry unique national identification numbers that function as universal identifiers in many countries. Redact the national ID number and any biometric chip indicator unless the recipient has a legitimate regulatory requirement for it.

Workflow for ID document redaction

Scan or photograph the document and export as JPEG. Drop it onto the OCR Redactor drop zone. After OCR extracts text, draw rectangles over the document number, MRZ block, photograph area, and any additional fields your analysis determined should be covered.3 Review the text panel to confirm those fields are replaced with block characters. Export the PNG. For repeated ID verification workflows, you can establish a consistent checklist of fields to redact for each document type. Consequently, the same workflow applies regardless of which country's document you are processing, because the OCR Redactor has no built-in document type recognition and treats all regions as image areas you mark manually.

Passport structure under ICAO 9303 and the MRZ encoding format

ICAO Doc 9303 defines the physical layout and data encoding standard for Machine Readable Travel Documents, including the TD3 format used by standard passport booklets. A TD3 passport MRZ occupies two lines of 44 characters each on the biographical data page. The first line encodes document type, issuing state, surname, and given names using uppercase Latin letters, digits, and a filler character. The second line encodes passport number, nationality, date of birth, sex, expiry date, and personal number, along with four check digits at positions 10, 20, 28, and 43 plus a composite check digit at position 44.4

Why the MRZ requires redaction even when visual fields are covered

The MRZ must be redacted for most disclosure purposes because any MRZ parser can recover all encoded identity fields from those two text lines in milliseconds, since the ICAO Doc 9303 series standardizes that zone as the machine-readable portion of the data page5. Covering only the photographed face and the printed document number in the visual zone is insufficient: a recipient with scanning software recovers the passport number, date of birth, personal number, and expiry date from the machine-readable zone. Your redaction rectangle must span the entire two-line MRZ block at the bottom of the biographical data page to prevent this recovery.

Minimum disclosure principles for remote verification contexts

Remote verification workflows often require only that a document is genuine and unexpired and that the holder's name matches a self-declaration. A contractor submitting a passport scan to prove right-to-work eligibility, for example, does not need to share their passport number, date of birth, personal number, or MRZ with the verifying party. Sharing only the minimum necessary fields is both a privacy principle under GDPR Article 5(1)(c) and a practical limit on damage if the received copy is later exposed in a data breach.1

Applying minimum disclosure in the OCR Redactor

In the OCR Redactor, draw rectangles over the passport number field in the visual data zone, the date of birth, the personal number field, the entire MRZ block, and the holder photograph. Leave the holder name, issuing country label, and expiry date visible if the verification workflow requires those fields. Export the PNG and confirm the remaining visible fields are limited to exactly what the receiving party needs for the stated verification purpose. This check takes under two minutes and produces a disclosure-ready image that carries no more identity information than the transaction requires.

The two-minute review step is worth doing before every share. CapyToolkit's OCR Redactor shows the covered words in the text panel as you draw, so you can confirm the remaining visible fields match the stated purpose rather than guessing from the image alone. Treat any field outside the recipient's need as something to cover, because a smaller disclosure footprint reduces the harm if the copy is later leaked or mishandled, so strip a passport's MRZ and photo first before sharing anything.

When to use this

Use this tool when you must share a copy of a passport or government ID with a landlord, employer, platform, or counterparty but want to minimize the personal information disclosed beyond what the specific purpose requires.

Examples

Sharing a passport for remote age verification on a platform

Cover the document number, photograph, and MRZ. Leave the date of birth visible. Export the PNG. The platform can verify your age without gaining access to the document's full identity payload.

Providing a driver's license copy to a landlord

Cover the license number and home address. Leave the name, date of birth, and photo visible for identity matching. Export the PNG. The landlord can confirm identity and age without receiving your license number.

Sources
  1. 1.

    European Union, "Regulation (EU) 2016/679 (General Data Protection Regulation)," eur-lex.europa.eu, 2016. https://eur-lex.europa.eu/eli/reg/2016/679/oj/

  2. 2.

    Identity Week, "US set to unveil new ePassport design this summer," identityweek.net, March 2022. https://identityweek.net/us-set-to-unveil-new-epassport-design-this-summer/

  3. 3.

    "tesseract-ocr/tesseract," GitHub, accessed June 2026. https://github.com/tesseract-ocr/tesseract

  4. 4.

    "Machine-readable passport," Wikipedia, accessed June 2026. https://en.wikipedia.org/wiki/Machine-readable_passport

  5. 5.

    ICAO, "Doc 9303 — Machine Readable Travel Documents," icao.int, accessed June 2026. https://www.icao.int/publications/doc-series/doc-9303

FAQ

In most jurisdictions, yes. There is no legal requirement to provide a complete unredacted copy for most private-sector verification purposes. CapyToolkit lets you disclose only the minimum information necessary for the verification purpose. Consult applicable law in your jurisdiction for any regulated context such as financial services KYC.

The MRZ is the two-line block of text at the bottom of the passport photo page containing name, document number, nationality, date of birth, expiry date, and check digits encoded in a standardized format. Software can read it instantly. Always cover the MRZ when sharing a passport scan unless the recipient has a specific legal requirement for it.

Yes. Draw a rectangle over the photograph area on the image canvas. The photo region is typically not detected as text by OCR but the visual redaction layer covers it regardless of OCR output. The exported PNG shows a black rectangle in the photo position.

For remote age or name verification purposes, yes, a redacted photo is less useful for facial matching. For in-person verification where the verifier can see you directly, a redacted photo copy is sufficient proof of documentation without exposing your biometric data to storage risks on the verifier's systems.

Tesseract can extract MRZ text from clean high-resolution scans, but MRZ fonts use OCR-B typeface designed for mechanical reading rather than human readers. Accuracy varies by scan quality. Regardless of OCR accuracy, draw the rectangle over the visible MRZ block to cover it in the exported image and text output.

FOIA Document Redaction: OCR Before Redacting Government Releases

FOIA documents often arrive as low-resolution scans without OCR. Government agencies routinely fulfill Freedom of Information Act requests by scanning paper records and producing the scans as flat image PDFs, with no searchable text layer. Before you can apply word-aware redaction to a FOIA release for re-publication or further disclosure, OCR must extract the text from the image.

The OCR Redactor applies Tesseract OCR to the scanned image and then allows you to draw redaction rectangles whose coverage removes matching words from the text export. FOIA releases that require privacy redaction before further sharing, such as removing third-party personal information from records about yourself or redacting exempt information from records you intend to publish, follow the same workflow used for any scanned document.

What this page covers

  • Non-agency-employee names named in the worked example for scanned emails
  • Informant references covered under exemption b(7)(C)
  • Witness names covered under exemption b(7)(C)
  • Personal contact information covered under exemption b(7)(C)

Opens the Offline OCR & Document Redactor with this section's checklist shown at the top of the tool.

Open in the tool →

Common FOIA exemptions requiring redaction

FOIA Exemption b(6) covers personal privacy: information that would constitute a clearly unwarranted invasion of personal privacy.1 Third-party names, addresses, phone numbers, and personnel file information in government records fall under this exemption. When re-publishing FOIA releases, redact third-party personal information to avoid creating the very privacy harm the exemption was designed to prevent. Exemption b(7)(D) covers confidential sources in law enforcement records, protecting the identities of individuals who provided information under an express or implied promise of confidentiality.

Law enforcement and informant identity protection

Exemption b(7)(C) covers law enforcement records where disclosure could reasonably be expected to constitute an unwarranted invasion of personal privacy, typically informant identities, witness names, and source contact details.1 Consequently, any FOIA release containing law enforcement records requires careful review for third-party identifiers before sharing. CapyToolkit processes each page locally, so sensitive government records stay on your device throughout the redaction workflow.

Handling low-resolution FOIA scans

FOIA releases often arrive as 150 DPI or even 100 DPI scans, well below the 300 DPI threshold for reliable OCR.2 Tesseract can extract text from lower-resolution scans with reduced accuracy on small fonts and faint characters. Run OCR on the raw FOIA scan and review the text panel for completeness before relying on it for word-aware redaction. For pages where OCR misses significant text, draw redaction rectangles based on visual inspection of the image rather than text panel coverage alone. Building on this, federal agencies are required under FOIA to provide reasonably segregable non-exempt portions, meaning the document should already have agency redaction marks, which you can use as guides for your own further redaction.3

Re-publishing FOIA releases with additional privacy redactions

Journalists, researchers, and advocacy organizations regularly re-publish FOIA releases after removing third-party personal information not covered by the agency's original redactions. The workflow is: drop each page image into the OCR Redactor, identify unredacted personal information using OCR output and visual review, draw rectangles over names and identifiers requiring privacy protection, and export the redacted PNG. For multi-page releases, process pages in batches and combine exports. Furthermore, maintain a log of your additional redactions and the exemption basis for each, as this documentation supports your editorial or legal justification if the redaction decisions are challenged. Review each exported page visually before publication to confirm that no visible identifier remains outside the rectangles you drew.

Appealing agency over-redaction and the Vaughn index process

When a government agency releases a FOIA document with redactions you believe are excessive, the administrative appeal process is the first step toward obtaining a less-redacted version. Most agencies allow 90 days from receipt of the initial response to file an administrative appeal.4 In the appeal, identify each specific redaction you challenge, state the exemption the agency cited, and explain why you believe the exemption does not apply or why the redacted portion is reasonably segregable from exempt content. The agency must respond to the appeal within 20 working days under the FOIA statute.

If the administrative appeal is denied, judicial review in federal district court is the next option. Courts reviewing FOIA redaction disputes require agencies to submit a Vaughn index: a document-by-document, redaction-by-redaction itemization of each withheld piece of information and the specific statutory exemption and factual basis for withholding it.5 Courts may also conduct in camera review of the unredacted documents to assess whether the exemption claims are valid. Vaughn indexing forces agencies to articulate their redaction rationale with specificity, which often reveals over-broad exemption claims that do not survive judicial scrutiny.

Using OCR to search FOIA releases before adding redactions

OCR on a FOIA scan creates a searchable text layer that enables systematic name and identifier searching before you begin adding your own redactions. Drop the page image into the OCR Redactor and review the text panel for third-party names, phone numbers, addresses, and identifiers you intend to cover in your re-publication. Searching the text panel for each target string confirms whether OCR detected it in that image before you rely on text-based identification for redaction placement. For names that OCR misread due to scan quality, the text panel search may not find the name, requiring visual identification from the image instead.

Managing multi-page FOIA releases in the OCR Redactor workflow

FOIA releases frequently run dozens or hundreds of pages. Processing each page individually through the OCR Redactor is the correct workflow, but the session-by-session approach requires organization to keep page exports ordered and complete. Before starting, number the source image files by page order using a consistent naming convention such as release-2026-01-001.jpg through release-2026-01-247.jpg. This naming makes it straightforward to verify completeness at the end: the exported PNG count should match the source image count.

Process pages that require redaction first, using the OCR text panel to identify pages with third-party identifiers. Pages that the agency already fully redacted (where the entire text layer is covered by agency black boxes) produce little useful OCR output and may not need your own additional redactions. Pages with mixed content, where agency redactions coexist with unredacted third-party identifiers, require the most careful review: the agency's black boxes appear as image regions in your redaction session, and you need to add your own rectangles for the unredacted identifiers visible elsewhere on the same page.

Combining exported pages into a publishable FOIA release document

After processing all pages, combine the exported PNGs into a single PDF for publication or distribution. Any local PDF assembly tool accepts a directory of PNG files and produces a multi-page PDF. For publications on the web, a searchable PDF is preferable to an image-only PDF, but adding a text layer requires a second OCR pass after combination, which reintroduces the same text extraction risk that the raster export was designed to eliminate. For most FOIA publication purposes, an image-only PDF with no text layer is the appropriate format: it ensures readers can view the document but cannot use automated tools to extract names or identifiers that survived the agency's and your combined redaction.

Keep the page order deterministic when you assemble the release. CapyToolkit's OCR Redactor redacts each page as a flat raster image, so the final PDF preserves the redaction on every page without a recoverable text layer that could be parsed for the names your rectangles removed. Naming the exported files in page sequence and combining them in that order avoids scrambled releases where a reader encounters page 40 before page 12, so scrub third-party names from FOIA releases page by page before publishing.

When to use this

Use this tool when preparing FOIA releases for publication, when adding privacy redactions to government records before sharing them with research partners, or when creating a redacted working copy of a FOIA release for reference during analysis.

Examples

Redacting third-party names from a FOIA-obtained government email release

Draw rectangles over all non-agency-employee names visible in the scanned emails. Export each page as PNG. Confirm in the text panel that the redacted names are replaced with block characters before exporting.

Preparing a law enforcement record FOIA release for publication

Identify all informant references, witness names, and personal contact information under b(7)(C). Draw rectangles over each. Add your redaction marks to any that the agency left visible in its original release.

Sources
  1. 1.

    Cornell Law School Legal Information Institute, "5 U.S. Code § 552. Public Access to Federal Records," law.cornell.edu, accessed June 2026. https://www.law.cornell.edu/uscode/text/5/552

  2. 2.

    National Archives and Records Administration, "1998 Digitization Guidelines," archives.gov, 1998. https://www.archives.gov/files/preservation/technical/guidelines-1998.pdf

  3. 3.

    U.S. Department of Justice, "FOIA Statute," foia.gov, accessed June 2026. https://www.foia.gov/foia-statute.html

  4. 4.

    U.S. Department of Justice, "Adjudicating Administrative Appeals Under the FOIA," justice.gov, accessed June 2026. https://www.justice.gov/oip/oip-guidance/Adjudicating%20Administrative%20Appeals%20Under%20the%20FOIA

  5. 5.

    U.S. Department of Justice, "FOIA Litigation Handbook," justice.gov, May 2021. https://www.justice.gov/oip/page/file/1399966/dl?inline=

FAQ

Yes. Drop the page image into CapyToolkit's OCR Redactor and check the text panel for the name you are searching for. If the text panel contains the name, draw a rectangle over that location on the image. For multiple instances across many pages, process each page individually and search each text panel for the target term.

Agency redaction marks appear as black boxes on the image. OCR will not extract text from those regions because the underlying content is already covered. Your additional redactions work the same way for any remaining unredacted identifiers on the same page.

FOIA does not impose obligations on requesters; it governs agency disclosure decisions. However, other laws, such as the Privacy Act for federal records and state privacy laws, may limit how you may re-publish information about identifiable private individuals. Consult applicable law for your specific jurisdiction and intended use.

Government agency scans vary widely. Clean 200 DPI scans of typed text produce good accuracy. Microfilm conversions, photocopied originals, and handwritten notes produce lower accuracy. Always supplement OCR-identified text with visual inspection when the OCR output appears incomplete or garbled.

The OCR Redactor draws opaque black rectangles. You cannot add text labels or annotation markers to indicate the basis for a redaction within the exported image. For annotated versions, export the PNG and add annotations in a separate image editor before final distribution.

FAQ

A box drawn as an annotation or shape sits on top of the page, and the text underneath often stays in the file, where copying, searching or a PDF parser can still read it. The OCR Redactor exports a flat image of the page, so there is no text layer under the black boxes to recover.

Yes, after turning each page into an image. The OCR Redactor accepts JPEG, PNG and WebP files up to 20 MB, so export each PDF page as an image or take a full-resolution screenshot of it, then redact the image and export it as PNG or PDF.

One page at a time. Export each page as an image, redact it, download the result, and combine the redacted pages into one PDF with a local PDF tool afterwards. Pages that carry identifiers, such as headers, signature pages and attachments, need the closest review.

No. CapyToolkit doesn't upload the documents you redact or send your scans to any server. The OCR engine runs in your browser, so the page image, the extracted text and the redacted export all stay on your device.

No. The exported PNG or PDF is drawn from the redacted canvas in your browser, not copied from your original file, so the scanner's metadata does not carry over. The original scan on your device keeps its metadata, so share only the redacted export.