FOIA Document Redaction: OCR and Redact Government Record Releases
FOIA documents often arrive as low-resolution scans without OCR. Government agencies routinely fulfill Freedom of Information Act requests by scanning paper records and producing the scans as flat image PDFs, with no searchable text layer. Before you can apply word-aware redaction to a FOIA release for re-publication or further disclosure, OCR must extract the text from the image.
The OCR Redactor applies Tesseract OCR to the scanned image and then allows you to draw redaction rectangles whose coverage removes matching words from the text export. FOIA releases that require privacy redaction before further sharing, such as removing third-party personal information from records about yourself or redacting exempt information from records you intend to publish, follow the same workflow used for any scanned document.
Common FOIA exemptions requiring redaction
FOIA Exemption b(6) covers personal privacy: information that would constitute a clearly unwarranted invasion of personal privacy.1 Third-party names, addresses, phone numbers, and personnel file information in government records fall under this exemption. When re-publishing FOIA releases, redact third-party personal information to avoid creating the very privacy harm the exemption was designed to prevent. Exemption b(7)(D) covers confidential sources in law enforcement records, protecting the identities of individuals who provided information under an express or implied promise of confidentiality.
Law enforcement and informant identity protection
Exemption b(7)(C) covers law enforcement records where disclosure could reasonably be expected to constitute an unwarranted invasion of personal privacy, typically informant identities, witness names, and source contact details.1 Consequently, any FOIA release containing law enforcement records requires careful review for third-party identifiers before sharing. CapyToolkit processes each page locally, so sensitive government records stay on your device throughout the redaction workflow.
Handling low-resolution FOIA scans
FOIA releases often arrive as 150 DPI or even 100 DPI scans, well below the 300 DPI threshold for reliable OCR.2 Tesseract can extract text from lower-resolution scans with reduced accuracy on small fonts and faint characters. Run OCR on the raw FOIA scan and review the text panel for completeness before relying on it for word-aware redaction. For pages where OCR misses significant text, draw redaction rectangles based on visual inspection of the image rather than text panel coverage alone. Building on this, federal agencies are required under FOIA to provide reasonably segregable non-exempt portions, meaning the document should already have agency redaction marks, which you can use as guides for your own further redaction.3
Re-publishing FOIA releases with additional privacy redactions
Journalists, researchers, and advocacy organizations regularly re-publish FOIA releases after removing third-party personal information not covered by the agency's original redactions. The workflow is: drop each page image into the OCR Redactor, identify unredacted personal information using OCR output and visual review, draw rectangles over names and identifiers requiring privacy protection, and export the redacted PNG. For multi-page releases, process pages in batches and combine exports. Furthermore, maintain a log of your additional redactions and the exemption basis for each, as this documentation supports your editorial or legal justification if the redaction decisions are challenged. Review each exported page visually before publication to confirm that no visible identifier remains outside the rectangles you drew.
Appealing agency over-redaction and the Vaughn index process
When a government agency releases a FOIA document with redactions you believe are excessive, the administrative appeal process is the first step toward obtaining a less-redacted version. Most agencies allow 90 days from receipt of the initial response to file an administrative appeal.4 In the appeal, identify each specific redaction you challenge, state the exemption the agency cited, and explain why you believe the exemption does not apply or why the redacted portion is reasonably segregable from exempt content. The agency must respond to the appeal within 20 working days under the FOIA statute.
If the administrative appeal is denied, judicial review in federal district court is the next option. Courts reviewing FOIA redaction disputes require agencies to submit a Vaughn index: a document-by-document, redaction-by-redaction itemization of each withheld piece of information and the specific statutory exemption and factual basis for withholding it.5 Courts may also conduct in camera review of the unredacted documents to assess whether the exemption claims are valid. Vaughn indexing forces agencies to articulate their redaction rationale with specificity, which often reveals over-broad exemption claims that do not survive judicial scrutiny.
Using OCR to search FOIA releases before adding redactions
OCR on a FOIA scan creates a searchable text layer that enables systematic name and identifier searching before you begin adding your own redactions. Drop the page image into the OCR Redactor and review the text panel for third-party names, phone numbers, addresses, and identifiers you intend to cover in your re-publication. Searching the text panel for each target string confirms whether OCR detected it in that image before you rely on text-based identification for redaction placement. For names that OCR misread due to scan quality, the text panel search may not find the name, requiring visual identification from the image instead.
Managing multi-page FOIA releases in the OCR Redactor workflow
FOIA releases frequently run dozens or hundreds of pages. Processing each page individually through the OCR Redactor is the correct workflow, but the session-by-session approach requires organization to keep page exports ordered and complete. Before starting, number the source image files by page order using a consistent naming convention such as release-2026-01-001.jpg through release-2026-01-247.jpg. This naming makes it straightforward to verify completeness at the end: the exported PNG count should match the source image count.
Process pages that require redaction first, using the OCR text panel to identify pages with third-party identifiers. Pages that the agency already fully redacted (where the entire text layer is covered by agency black boxes) produce little useful OCR output and may not need your own additional redactions. Pages with mixed content, where agency redactions coexist with unredacted third-party identifiers, require the most careful review: the agency's black boxes appear as image regions in your redaction session, and you need to add your own rectangles for the unredacted identifiers visible elsewhere on the same page.
Combining exported pages into a publishable FOIA release document
After processing all pages, combine the exported PNGs into a single PDF for publication or distribution. Any local PDF assembly tool accepts a directory of PNG files and produces a multi-page PDF. For publications on the web, a searchable PDF is preferable to an image-only PDF, but adding a text layer requires a second OCR pass after combination, which reintroduces the same text extraction risk that the raster export was designed to eliminate. For most FOIA publication purposes, an image-only PDF with no text layer is the appropriate format: it ensures readers can view the document but cannot use automated tools to extract names or identifiers that survived the agency's and your combined redaction.
Keep the page order deterministic when you assemble the release. CapyToolkit's OCR Redactor redacts each page as a flat raster image, so the final PDF preserves the redaction on every page without a recoverable text layer that could be parsed for the names your rectangles removed. Naming the exported files in page sequence and combining them in that order avoids scrambled releases where a reader encounters page 40 before page 12, so scrub third-party names from FOIA releases page by page before publishing.
When to use this
Use this tool when preparing FOIA releases for publication, when adding privacy redactions to government records before sharing them with research partners, or when creating a redacted working copy of a FOIA release for reference during analysis.
Examples
Redacting third-party names from a FOIA-obtained government email release
Draw rectangles over all non-agency-employee names visible in the scanned emails. Export each page as PNG. Confirm in the text panel that the redacted names are replaced with block characters before exporting.
Preparing a law enforcement record FOIA release for publication
Identify all informant references, witness names, and personal contact information under b(7)(C). Draw rectangles over each. Add your redaction marks to any that the agency left visible in its original release.
- 1.
Cornell Law School Legal Information Institute, "5 U.S. Code § 552. Public Access to Federal Records," law.cornell.edu, accessed June 2026. https://www.law.cornell.edu/uscode/text/5/552
- 2.
National Archives and Records Administration, "1998 Digitization Guidelines," archives.gov, 1998. https://www.archives.gov/files/preservation/technical/guidelines-1998.pdf
- 3.
U.S. Department of Justice, "FOIA Statute," foia.gov, accessed June 2026. https://www.foia.gov/foia-statute.html
- 4.
U.S. Department of Justice, "Adjudicating Administrative Appeals Under the FOIA," justice.gov, accessed June 2026. https://www.justice.gov/oip/oip-guidance/Adjudicating%20Administrative%20Appeals%20Under%20the%20FOIA
- 5.
U.S. Department of Justice, "FOIA Litigation Handbook," justice.gov, May 2021. https://www.justice.gov/oip/page/file/1399966/dl?inline=