Scanner Settings for OCR and Redaction

Extract text from images and redact sensitive regions locally. Tesseract WASM runs in your browser — no uploads, no server, no account needed.

Scanner Settings for OCR and Redaction

OCR accuracy is mostly decided before the page reaches any OCR engine. A scan at too low a resolution, squeezed by heavy JPEG compression or tinted by colour noise gives Tesseract broken letter shapes to work with, and every misread word is a word the text panel cannot help you find and redact. The good news is that the settings that fix this are the same on almost every scanner and scanning app.

Each device then adds its own obstacle: a default that routes scans to the cloud, a driver that breaks after an update, or a compression level you cannot see in the app. The section below covers the shared settings once. The entries cover the Fujitsu ScanSnap iX1600, the Canon imageFORMULA R40, the Epson WorkForce ES-400 II, Adobe Scan and the iPhone's built-in document scanner.

Scan settings for OCR

  • 300 DPI for standard printed text
  • 400 to 600 DPI
  • grayscale unless colour carries meaning
  • PNG, or JPEG at the highest quality setting
  • 20 MB per image

Check the extracted text after OCR: garbled words in a clean-looking scan usually point to compression or resolution, not to the page.

Opens the Offline OCR & Document Redactor with this page's reference values shown at the top of the tool.

Open in the tool →

Settings that matter on every scanner

Three settings do most of the work: resolution, colour mode and file format, plus one choice about where the finished file goes. Every device on this page exposes them in some form, whether in a desktop scan profile or in a phone app's export options, so getting them right once turns the per-device quirks below into small adjustments rather than repeat scans.

Resolution

Tesseract's own documentation says it works best on images with a resolution of at least 300 DPI.1 That is the right setting for ordinary printed text at 10 to 12 points, where each letter is tall enough in pixels for the engine to tell similar shapes such as l, 1 and I apart.

Go to 400 or 600 DPI for footnotes, small print in contracts and dense financial tables, and expect larger files in return. Higher is not always better, because a 600 DPI colour scan of a letter-size page can approach or pass the 20 MB limit the OCR Redactor accepts, and the extra detail adds nothing for normal body text.

Colour mode and format

Grayscale keeps the letter edges and drops the colour noise that paper texture and coloured backgrounds add, which is why it is the safe default for text documents. Use colour only when it matters for the redaction, such as highlighted fields, coloured stamps or signatures in blue ink that you need to find and cover.

For the file format, PNG keeps every pixel as scanned, while JPEG trades detail for size. If a device only offers JPEG, pick its highest quality setting, because the blocky compression artefacts at lower settings blur exactly the thin strokes OCR depends on, and a misread word is a word you cannot find in the text panel.

Where the file goes

Many scanners and scanning apps send finished scans to a cloud service by default. For documents you are about to redact, that defeats the point, so set the destination to a local folder before the first scan and move files to your computer by USB, a local network share or AirDrop. The unredacted original is the most sensitive version of the document, so it should be the one that travels least.

Sources
  1. 1.

    Tesseract OCR, "Improving the quality of the output," tesseract-ocr.github.io, accessed October 2026. https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html

Fujitsu ScanSnap iX1600 OCR and Redaction Guide

ScanSnap iX1600 needs one setting change for OCR. Switch the quality mode from "Auto" to "Best" in ScanSnap Home before scanning any document intended for text extraction. Auto mode applies aggressive JPEG compression that introduces block artifacts at text edges, causing Tesseract to misread characters along word boundaries1. At the Best setting, the scanner outputs a higher-quality JPEG at 300 DPI, preserving the sharp ink-to-paper contrast that OCR depends on2.

ScanSnap Home routes finished scans to Adobe Document Cloud or ScanSnap Cloud by default. Disable cloud routing in the destination settings and choose a local folder instead. Keeping scans on your machine is a prerequisite for feeding them into a browser-based tool like the OCR Redactor without any upload occurring. After those two changes, the iX1600 produces output that Tesseract processes with high accuracy on standard printed documents.

Run this check yourself in the Offline OCR & Document Redactor.

Open in the tool →

Specifications3

ADF Capacity50 sheets
Scanning speed40 ppm simplex / 80 ipm duplex
Max resolution600 DPI
ConnectionUSB 3.2 Gen 1 or Wi-Fi 5
SoftwareScanSnap Home

ScanSnap Home quality and colour settings for OCR

Inside ScanSnap Home, open the scanning profile and set Quality to Best and Color mode to Auto Color Detection. Best quality scans at 300 DPI instead of the compressed 150 DPI used by Auto mode3. Scanning at 300 DPI ensures that letter-height characters contain enough pixel rows for Tesseract to distinguish similar glyphs such as "l", "1", and "I"1. Duplex scanning captures both sides of a sheet in one pass and works well for double-sided contracts. Confirm the output format is JPEG rather than PDF before dragging the file into the OCR Redactor, since the tool accepts image formats only and does not parse PDF containers. Consequently, JPEG at Best quality is the correct output format for this workflow.

Double feeds and glossy paper on the iX1600

Fujitsu iX1600 firmware versions below 1.6 produced intermittent double-feed errors when scanning coated paper. Update the firmware through ScanSnap Home under Settings > Version Information > Check for Updates. A separate documented issue involves the ADF picking two sheets simultaneously without triggering the ultrasonic multi-feed sensor, which calibrates for standard 80 g/m² paper. Glossy pages fool the sensor; place one sheet at a time for glossy printouts.

Roller maintenance and scan alignment

The roller kit requires replacement after approximately 200,000 sheets4. Worn rollers cause misfeeds and skewed scans that produce trapezoidal text blocks, degrading OCR accuracy compared to a straight horizontal scan. CapyToolkit's OCR Redactor processes the output from any scanner that saves locally, so keeping the iX1600 rollers in good condition directly impacts how accurately Tesseract reads your scanned text. Fujitsu sells the roller kit as a replacement part through its business supply portal, and the iX1600 displays a replacement warning when the sheet count approaches the maintenance threshold.

Feeding ScanSnap output into the OCR Redactor

After ScanSnap Home saves the finished scan to your local folder, drag the JPEG file onto the OCR Redactor drop zone. Tesseract starts processing immediately and returns extracted text alongside the original image. Your redaction rectangles cover the sensitive regions you draw over. Because ScanSnap Home defaults to JPEG compression even at Best quality, check the extracted text for garbled characters near edges where JPEG blocks can blur ink strokes. Rescanning as PNG using the ScanSnap Home profile editor eliminates compression entirely and improves accuracy on documents with very small fonts or dense tables5. Yet for most standard printed documents at 300 DPI, JPEG Best quality produces reliable OCR output.

Configuring ScanSnap Home scan profiles for different document types

Open ScanSnap Home and create a named profile for each document category you regularly process. Each profile stores your resolution, color mode, and output destination as a saved configuration you can select directly from the iX1600 touchscreen. Creating separate profiles for standard forms at 300 DPI Grayscale and small-font ID documents at 400 DPI Color prevents manual reconfiguration before each scan batch.

Optimizing profile settings for OCR accuracy

Resolution is the most consequential profile setting for downstream OCR quality. Documents with body text at 11pt or larger scan reliably at 300 DPI in Grayscale mode. Forms with 8pt or smaller labels and identifier fields benefit from 400 DPI, which Tesseract processes without additional preprocessing. Avoid the Auto Color setting for documents destined for OCR: it introduces JPEG compression artifacts into grayscale regions of mixed-color pages, which can degrade edge contrast on thin letterforms and produce character substitution errors in the extracted text.

Using the iX1600 touchscreen and Wi-Fi 5 for scanning without a host computer

The iX1600 includes a five-inch color touchscreen and dual-band Wi-Fi 5 (802.11ac) that enable cloud-connected scanning without an active USB cable. From the touchscreen, select a saved profile and scan directly to a ScanSnap Cloud destination such as Google Drive or Dropbox, or to a local network folder configured in ScanSnap Home. This works when you need to scan at a desk that lacks an attached workstation.

Routing scans to a local folder without cloud transit

Connecting to a local network folder via ScanSnap Cloud requires the ScanSnap Home server component running on another computer on your Wi-Fi network. That computer does not need to be the workstation you use for redaction; it only needs ScanSnap Home installed and a designated shared folder. Once the scan arrives in the shared folder, move it to the computer running the OCR Redactor and drag it onto the drop zone as normal, following the local-only ScanSnap iX1600 redaction path. This approach keeps all document data on your local network without routing files through a cloud storage intermediary.

For recurring batch work, configure ScanSnap Home to deliver every finished scan straight to the shared folder so nothing is staged in a cloud inbox first. The OCR Redactor reads only the local file you drag into the drop zone, so the document never traverses an external network during the redaction step. This local-only path is the reason the iX1600 fits workflows that must keep source documents on premises.

Sources
  1. 1.

    Tesseract OCR, "Improving the quality of the output," tesseract-ocr.github.io, accessed June 2026. https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html

  2. 2.

    Ricoh, "ScanSnap iX1600," pfu.ricoh.com, accessed June 2026. https://www.pfu.ricoh.com/global/scanners/scansnap/ix1600/

  3. 3.

    Ricoh, "ScanSnap iX1600," pfu-asia.ricoh.com, accessed June 2026. https://www.pfu-asia.ricoh.com/sg/product/scansnap-ix1600/

  4. 4.

    Ricoh, "Replacing the Roller Set," pfu.ricoh.com, accessed June 2026. https://www.pfu.ricoh.com/imaging/downloads/manual/ss_webhelp/en/help/webhelp/topic/ma_consumable_rollerset.html

  5. 5.

    Tesseract OCR, "PNG vs JPEG for OCR accuracy," github.com, accessed June 2026. https://github.com/tesseract-ocr/tesseract/issues/1895

FAQ

ScanSnap Home routes scans to a cloud destination by default. Open the profile editor, click the destination tab, and switch from any cloud option to Computer. After this change, finished scans land in the local folder you specify, ready to drag into the OCR Redactor without an upload step.

300 DPI is the standard minimum for printed text OCR. Set Quality to Best in ScanSnap Home, which outputs at 300 DPI. Handwritten documents may benefit from 400 DPI or higher. The iX1600 supports up to 600 DPI, which covers any practical OCR need.

The ultrasonic multi-feed sensor calibrates for 80 g/m² paper. Glossy and coated stock confuses it. Disable the multi-feed sensor in the profile settings and feed glossy pages one at a time manually. Re-enable the sensor for normal paper to prevent actual double-feeds from going undetected.

Yes. Any scanner that saves images to a folder works with the OCR Redactor. Set the iX1600 to save locally as JPEG or PNG, then drag the file from that folder into the drop zone. CapyToolkit does not require any proprietary software on the scanning side. ScanSnap Home is not required after the scan is saved to your local drive.

Diagonal lines on ADF scans usually come from debris on the scanning glass strip inside the document path. Open the ADF cover, locate the narrow glass strip near the rollers, and wipe it with the included cleaning cloth. Even a small particle of dust creates a consistent dark line through every page that passes over it.

Canon imageFORMULA R40 OCR and Redaction Guide

Canon R40 connects via USB 2.0 with no wireless option1. On Windows 11, install the driver from Canon's R40 Software Installer for Windows, which bundles the scanner driver, CaptureOnTouch and the user manual and lists Windows 11 among its supported systems2. If a scanning app reports the R40 driver as damaged or missing after a Windows update, reinstalling that package with the scanner disconnected is the first fix to try.

The R40 bundles ReadIris Pro for OCR and document management. For large documents, the OCR Redactor's Tesseract WASM engine provides a more reliable alternative with no server connection required. Set the CaptureOnTouch scan job to save JPEG files at 300 DPI to a local folder, then drop each file into the OCR Redactor drop zone.

Run this check yourself in the Offline OCR & Document Redactor.

Open in the tool →

Specifications3

ADF Capacity60 sheets
Scanning speed40 ppm simplex / 80 ipm duplex
Max resolution600 DPI
ConnectionUSB 2.0 (no Wi-Fi)
SoftwareCanon CaptureOnTouch, ReadIris Pro (bundled)

CaptureOnTouch scan job settings for OCR

Inside Canon CaptureOnTouch, create a scan job with resolution 300 DPI, color mode Grayscale, and output format JPEG. Grayscale at 300 DPI preserves tonal contrast better than Black & White for documents with faint printed text or carbon-copy forms. Setting the brightness to +5 in the scan job lightens background paper yellowing on older documents without washing out ink.3

ADF batch scanning and per-job output folders

Because the R40 has a 60-sheet ADF capacity, scanning an entire contract in one pass produces a folder of numbered JPEG files. Drag each page separately into the OCR Redactor, or process only the most sensitive pages. CaptureOnTouch allows per-job output folder configuration, keeping scan batches organized by project. CapyToolkit's OCR Redactor accepts each JPEG individually, so you can prioritize the pages that contain sensitive information first without processing the entire batch.

Driver errors after Windows updates and ReadIris crashes

A driver error that appears right after a Windows update usually means the scanning app can no longer load the R40's TWAIN driver. Canon's R40 Software Installer for Windows bundles that driver with CaptureOnTouch and lists Windows 11 as supported2, so install it rather than the generic TWAIN driver Windows Update provides. Scanning at the R40's native 600 DPI setting produces noticeably sharper character edges in the captured image, which helps Tesseract distinguish fine letterforms and small printed text that would blur together at lower resolutions4. Furthermore, ReadIris crashes on large files because it loads the entire document image set into RAM simultaneously. The OCR Redactor processes one image at a time in a browser WebWorker, avoiding this memory issue entirely. The R40 needs its own AC power connection, so an unpowered or shared USB hub can leave the scanner without enough current to initialize reliably.

Feeding R40 output into the OCR Redactor

After CaptureOnTouch completes the scan job, open the output folder and select the JPEG file for the page you want to redact. Drag it onto the OCR Redactor drop zone. Tesseract begins extracting text immediately. Your drawn redaction boxes cover the marked regions, and the text panel simultaneously redacts the corresponding words. Exporting the result produces a PNG with opaque black rectangles over the redacted zones, with no underlying text layer. Conversely, PDF export from the tool produces a raster PDF with no searchable text, removing the hidden layer vulnerability common in standard PDF redaction software5. This makes the R40 and OCR Redactor combination suitable for preparing documents before sharing with external parties.

Organizing CaptureOnTouch scan jobs by project and output naming

CaptureOnTouch stores scan jobs as named profiles in a sidebar list. Creating one profile per document type or client project keeps your workflow organized when you handle several document categories in a single session. Open CaptureOnTouch, select New Scan Job in the sidebar, and configure scanning parameters before assigning a distinctive name to the profile.

Using output naming tokens for searchable filenames

The output file naming section in each scan job supports tokens for date, time, and an auto-incrementing counter. A pattern such as ProjectCode_YYYYMMDD_### produces files that sort correctly in Explorer and carry a readable date stamp per session. Combining this with a project-specific output folder path prevents files from different clients from intermixing in the same directory. These naming conventions simplify the handoff to the OCR Redactor: scan a batch, select the correct file by name, and drop it onto the drop zone without sorting through a folder of generic numbered filenames.

CaptureOnTouch Lite for portable scanning without installation

CaptureOnTouch Lite ships pre-installed on the R40 scanner itself and runs directly from the scanner as a USB storage device, without requiring a software installation on the host computer. Plug the R40 into any Windows or macOS computer, navigate to the CaptureOnTouch Lite application on the scanner's storage partition, and run it to begin scanning. Output files land in a folder you specify on the host computer or in the scanner's internal storage.

Moving Lite output to the OCR Redactor

After scanning with CaptureOnTouch Lite, locate the output JPEG in the configured folder or the scanner's storage partition. Copy it to the local computer before opening the OCR Redactor in a browser, since the tool reads files through drag and drop from the local file system. For recurring portable workflows, setting the output folder to a shared network path removes the manual copy step. CaptureOnTouch Lite does not support the full set of naming tokens available in the installed version, so adopt a brief manual naming convention for files produced during portable sessions.

Keeping the portable scan on local storage preserves the privacy advantage of the redaction step. CapyToolkit's OCR Redactor never uploads the image, so the document stays on the machine from the moment CaptureOnTouch Lite writes the JPEG until you export the redacted PNG. For field work away from the main office, this local-first path means sensitive contracts can be scanned, redacted, and saved without ever touching a cloud account, using the CaptureOnTouch Lite portable redaction route.

Sources
  1. 1.

    Canon Europe, "Canon imageFORMULA R40 - Document Scanners," canon-europe.com, accessed October 2026. https://www.canon-europe.com/business/products/scanners/document-scanners/imageformula-r40/

  2. 2.

    Canon Asia, "R40 Software Installer for Windows," asia.canon, accessed October 2026. https://asia.canon/en/support/0101056207?model=R40

  3. 3.

    Canon, "WorkForce ES-400 II / imageFORMULA R40," canon-europe.com, accessed June 2026. https://www.canon-europe.com/business/products/scanners/document-scanners/imageformula-r40/specifications/

  4. 4.

    Tesseract OCR Documentation, "Improving the Quality of the Output," tesseract-ocr.github.io, accessed June 2026. https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html

  5. 5.

    Adobe, "Redact and Sanitize PDFs in Acrobat Pro," helpx.adobe.com, updated September 2025. https://helpx.adobe.com/acrobat/desktop/protect-documents/redact-pdfs/redacting-sanitizing.html

FAQ

Windows installed an incompatible generic TWAIN driver over Canon's. Disconnect the scanner, go to Canon's support site, download the Windows 11 CaptureOnTouch plus ISIS/TWAIN Driver package, run the installer, then reconnect. Do not use Windows Update to install the TWAIN driver; always download it from Canon's official support page.

CapyToolkit's OCR Redactor processes one image at a time in a browser WebWorker and does not load an entire document into RAM at once. Drop each scanned page individually into the tool. There is no file count limit per session; process pages in sequence without any memory pressure from large batches.

Yes, for basic scanning. Windows 10 and 11 include a built-in scanner interface accessible from the Windows Scan app in the Microsoft Store. Set the resolution to 300 DPI and format to JPEG. CaptureOnTouch provides more control over batch naming and output folders but is not required for basic single-page scans.

Yes. The R40 is a duplex scanner rated at 80 ipm in duplex mode. Enable duplex scanning in the CaptureOnTouch job settings. Each sheet produces two image files. Feed each image into the OCR Redactor separately if you need to redact content on both sides.

The R40 runs on AC mains power and consumes up to 22 W while scanning, so its power adapter must be plugged into a working outlet. A powered USB hub alone will not bring the scanner up if the adapter is unplugged or on a circuit that is out of service. Connect the AC adapter first, wait for the scanner to settle, and only then attach the USB cable to the computer.

Epson WorkForce ES-400 II OCR and Redaction Guide

Epson ES-400 II scans at 35 pages per minute via ADF1. Connecting via USB 3.0 rather than USB 2.0 speeds up image transfer significantly during large batch scans1. Epson ScanSmart software manages the connection, but a documented issue on some Windows 10 and 11 setups causes the Scanner Settings window to appear blank immediately after opening2.

Updating ScanSmart to the latest version from Epson's support page resolves the blank settings window in most cases. Alternatively, reinstalling the WIA driver separately from the ScanSmart installer provides a fallback connection path. Once configured correctly, the scanner delivers clean 300 DPI text scans that feed into the OCR Redactor without preprocessing.

Run this check yourself in the Offline OCR & Document Redactor.

Open in the tool →

Specifications3

ADF Capacity50 sheets
Scanning speed35 ppm simplex / 70 ipm duplex
Max resolution600 DPI
ConnectionUSB 3.0 (USB 2.0 compatible)
SoftwareEpson ScanSmart

ScanSmart resolution and colour mode for OCR

Set the resolution to 300 DPI and color mode to Black & White for text-only documents in Epson ScanSmart3. Grayscale and color modes produce larger files with no OCR benefit for standard black ink on white paper. Set the paper size to Auto to let ScanSmart detect page boundaries correctly. Scanning at 300 DPI Black & White produces compact images with high contrast, which Tesseract processes faster than color scans4. The ES-400 II supports up to 600 DPI optical resolution, but 300 DPI is the practical sweet spot for OCR because higher resolutions produce diminishing accuracy returns while significantly increasing file size and processing time.

Choosing the right color mode for your document type

For documents with colored stamps, highlights, or watermarks that OCR needs to read, switch to Grayscale; this preserves tonal variation without the three-channel overhead of full color. Consequently, Black & White is the default and Grayscale is the exception. When you plan to redact colored elements alongside printed text, scanning in Grayscale ensures the redaction rectangles cover the correct regions without losing tonal detail that distinguishes stamps from body text. For forms where the only color is a light blue or gray background grid, Black & White remains the better choice because the grid binarizes cleanly and does not interfere with character recognition.

Blank settings window, streaks and banding on the ES-400 II

The blank Scanner Settings window appears when ScanSmart cannot communicate with the TWAIN driver2. Reinstalling ScanSmart from Epson's support site while the scanner is disconnected, then reconnecting after installation, resolves this in most cases. ADF scans with faint vertical lines running the full page height indicate debris on the scanning glass strip inside the roller path. Wipe the glass with a dry microfiber cloth. Furthermore, horizontal bands in the output point to roller wear; replace the roller kit when band artifacts appear consistently across multiple sheets. USB 2.0 connections also introduce noticeable transfer delays on large batches because the interface saturates at lower throughput than USB 3.0.

Feeding ES-400 II output into the OCR Redactor

After ScanSmart saves the scan, locate the JPEG or PNG file in your output folder and drag it onto the OCR Redactor drop zone. ScanSmart defaults to saving multi-page scans as a PDF; configure the output format to JPEG or PNG in the file type settings before scanning. Because the OCR Redactor accepts single image files, each page must be a separate image file. For a 20-page contract, scan each page individually or configure ScanSmart to save each page as a separate JPEG. Building on this, export at 300 DPI color only when documents contain colored stamps or handwriting that the redaction must cover alongside printed text.

Duplex scanning on the ES-400 II and page ordering for multi-page redaction

Duplex scanning captures both sides of a sheet in a single ADF pass, which the ES-400 II handles at 70 images per minute. Epson ScanSmart saves duplex results either as a single multi-page file or as separate files per image, depending on the output format setting. For the OCR Redactor workflow, set the output to save each image as a separate JPEG: front and back of each page become two numbered files. This naming convention keeps the page ordering clear when you process a multi-page document.

ScanSmart's default simplex-first duplex ordering saves images in the sequence: front of page 1, back of page 1, front of page 2, back of page 2. Confirm this matches your expectations before processing a long document, because some ADF drivers use an alternate ordering (all fronts first, then all backs) that requires reordering before combining into a final PDF. For the OCR Redactor, process each numbered JPEG individually, export the redacted PNG, and combine in the correct page order when assembling the final multi-page document.

Using ScanSmart profiles to create a reusable OCR-ready scan configuration

ScanSmart supports named scan profiles that save all settings: resolution, color mode, paper size, file format, and destination folder. Creating a dedicated "OCR Redactor" profile set to 300 DPI, Black & White, JPEG output, and a specific local folder eliminates the need to reconfigure settings before each redaction session. Save the profile once, select it at the start of each scan session, and all output files land in the correct folder at the correct specification automatically. For organizations processing multiple document types, create separate profiles for text-only documents (300 DPI Black & White) and documents with colored elements requiring redaction (300 DPI Color), switching between profiles based on document content before scanning.

Connecting the ES-400 II to network-accessible workstations for shared use

The ES-400 II connects via USB 3.0 only, with no built-in Wi-Fi or Ethernet. In shared office environments, placing the scanner on a dedicated workstation and using Windows Shared Folders or macOS File Sharing to expose the scan output folder allows other network users to access completed scans without physically connecting to the scanner. The scanning workstation runs ScanSmart and saves files directly to the shared folder; remote users retrieve the files via the network share and drop them into the OCR Redactor on their own machines.

This architecture separates the scanning hardware from the redaction workstation, the approach behind the Epson scanner and redaction workstation split, which matters for compliance workflows where the person performing redaction should not be the same person who scanned the document. It also avoids the security exposure of running a browser-based OCR tool on the same machine connected to a scanner that may have cached prior scans. Configuring the shared folder with read-only access for users other than the designated scanner operator prevents accidental modification or deletion of original scan files before redaction is complete.

Maintaining the ES-400 II roller kit for consistent scan quality

Epson rates the ES-400 II roller assembly kit for approximately 200,000 cycles before replacement is recommended5. Worn rollers produce two visible failure modes: inconsistent feed speed that causes horizontal banding in the scan output, and intermittent multi-feeds where two sheets advance simultaneously without triggering the ultrasonic sensor. Both failures degrade OCR quality: banding distorts character shapes, and multi-feeds produce combined images of two document pages that Tesseract processes as a single image with garbled layout. Order the ES-400 II Roller Assembly Kit (part number B12B819671) from Epson's accessories page when the sheet count approaches 200,000 or when banding artifacts appear consistently across separate scan sessions.

Replacing the roller kit on schedule protects the accuracy of every downstream OCR job. The OCR Redactor relies on clean, evenly fed scans to place bounding boxes correctly, and banding from worn rollers shifts character shapes enough to lower recognition rates. CapyToolkit's OCR Redactor ingests the resulting JPEG or PNG files locally, so the maintenance habit on the scanner side is what keeps the text layer reliable.

Sources
  1. 1.

    Epson, "WorkForce ES-400 II Duplex Desktop Document Scanner," epson.com, accessed June 2026. https://epson.com/For-Home/Scanners/Document-Scanners/WorkForce-ES-400-II-Duplex-Desktop-Document-Scanner/p/B11B261201

  2. 2.

    Microsoft, "Epson ScanSmart software will not launch after Windows update," learn.microsoft.com, accessed June 2026. https://learn.microsoft.com/en-us/answers/questions/4120210/epson-scansmart-software-will-not-launch-after-win

  3. 3.

    Tesseract OCR, "Improving the quality of the output," tesseract-ocr.github.io, accessed June 2026. https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html

  4. 4.

    Epson, "ES-400 II/ES-500W II User's Guide," files.support.epson.com, 2024, pp. 1–2. https://files.support.epson.com/docid/cpd5/cpd59603.pdf

  5. 5.

    Epson, "Roller Assembly Kit B12B819671," epson.com, accessed June 2026. https://epson.com/Accessories/Scanner-Accessories/Roller-Assembly-Kit-B12B819671/p/B12B819671

FAQ

This is a documented bug in some ScanSmart versions on Windows 10 and 11. Download the latest ScanSmart installer from Epson's support site, uninstall the current version with the scanner disconnected, then reinstall and reconnect. If the blank window persists, download and install the WIA driver package separately from the ScanSmart installer.

300 DPI in Black & White mode is sufficient for printed documents. Epson ScanSmart defaults to 200 DPI, which can miss fine serifs on small fonts. Change it to 300 DPI manually in the scanning profile. Handwritten documents benefit from 400 DPI.

USB 2.0 sustained throughput is lower than USB 3.0. The speed difference becomes noticeable in large 50-sheet batch scans at high resolution. Connect to a USB 3.0 port when scanning full batches. For single pages, the difference is minimal.

Vertical lines running the full page height come from debris on the scanning glass strip inside the document path. Open the ADF cover and locate the thin glass strip near the rollers. Wipe it with a dry microfiber cloth. A single dust particle on that strip creates a consistent line through every page in the batch.

Yes. CapyToolkit's OCR Redactor accepts any JPEG or PNG file up to 20 MB regardless of which scanner software produced it. Use the Windows built-in scanner interface, set the format to JPEG or PNG, and drag the saved file into the tool. ScanSmart is not a requirement.

Adobe Scan: Disabling Cloud Sync and Redacting Locally

Adobe Scan uploads every document to Adobe Document Cloud by default using the Adobe Sensei engine for server-side OCR1. Signing out of your Adobe ID in the app prevents scans from syncing to the cloud, and disabling cellular data for Adobe Scan in your device settings blocks all network access. Without those precautions, every page you scan transmits to Adobe's servers before you can process it locally.

Adobe Scan produces solid perspective correction on mobile photos, handling moderate tilt from handheld capture. Its compression pipeline outputs JPEG files at a moderate quality level even when the source photo is high resolution. Consequently, exported files often contain visible JPEG block artifacts around fine text strokes, which can reduce Tesseract's character recognition accuracy on fonts below 10 points2. After blocking network access, local JPEG exports work directly with the OCR Redactor.

Run this check yourself in the Offline OCR & Document Redactor.

Open in the tool →

Specifications3

PlatformiOS and Android (free)
Output formatJPEG or PDF
OCRServer-side by default (Adobe Sensei)
Installs100M+ (2026)
ExportShare as JPEG or PDF via Files/Photos app

Disabling cloud uploads before scanning

Inside Adobe Scan, tap your profile avatar in the upper right, then Settings. Sign out of your Adobe ID and disable cellular data for the app in your device system settings. This prevents the app from transmitting the scan to Adobe servers immediately after capture. Also turn off "Save originals to Camera Roll" if your camera roll syncs to iCloud or Google Photos, as that creates a second upload path through a different cloud service4.

Exporting local scans for desktop redaction

After disabling cloud sync, finished scans remain local on the device. Export them by tapping Share and choosing Save Image, which writes the JPEG to your device photo library. You can then transfer the file to your computer and drop it into the OCR Redactor. CapyToolkit processes the file entirely in the browser, so even after transfer, no document data passes through any server during the OCR and redaction workflow.

The local export step is what makes the whole workflow private. CapyToolkit's OCR Redactor processes the JPEG entirely in the browser, so the only movement the document makes is from your phone to your own computer and then into the tool's drop zone. Because nothing is sent to a server, you can keep the original scan on the device until you confirm the redacted export looks correct before deleting the source.

Common issues with scan quality for OCR

Adobe Scan handles moderate tilt well during perspective correction, but the output JPEG quality level is lower than a dedicated ADF scanner. Scans taken under fluorescent lighting with shadows across the page produce uneven contrast that Tesseract struggles to segment correctly. Photograph the document under natural light or direct overhead lighting with no shadows, holding the phone directly above the page. Building on this, Adobe Scan applies automatic color balance that can wash out gray text on colored paper backgrounds. For documents with colored backgrounds, a dedicated scanner with adjustable gamma settings produces more consistent output than a mobile photo app.

Feeding Adobe Scan output into the OCR Redactor

Export the scan from Adobe Scan as a JPEG by tapping Share > Save Image. Transfer the file to your computer via USB, AirDrop, or any local transfer method. Drop the JPEG onto the OCR Redactor drop zone. Tesseract extracts text from the image, and you draw redaction rectangles over sensitive regions. Because Adobe Scan's JPEG compression can soften fine characters, check the OCR result panel for missing words near highly compressed regions2. Rescanning the same document with better lighting often recovers the missed characters. Furthermore, for multi-page documents, export each page individually and process them one at a time in the tool.

Improving Adobe Scan output quality for OCR accuracy

Adobe Scan applies automatic perspective correction and brightness normalization during capture, but these processes have limits. Perspective correction handles angles up to approximately 30 degrees reliably; beyond that, character shapes distort enough to reduce OCR accuracy on small fonts. Hold the phone directly above the document surface, parallel to the page, to stay within the correction range. For table-heavy documents with thin grid lines, perspective distortion causes individual cell contents to overlap adjacent cells after correction, which Tesseract's layout analysis then segments incorrectly. When a capture cannot be repeated, correcting the geometry after export is the fix: rotate the flattened page until the text lines run horizontal before running OCR, which is the same remedy Tesseract documents for skewed scans. ImageMagick's command line does this in one step with its rotate operator5.

Lighting is the most controllable quality factor for mobile capture. Adobe Scan's automatic tone mapping adjusts for overall exposure, but it cannot eliminate directional shadows that create brightness gradients across the text region. Place the document on a flat surface under overhead lighting, turn off any desk lamp that creates a shadow from one side, and ensure no window light hits the document at an angle. Under even lighting, the brightness gradient disappears and Adobe Scan's binarization step produces a clean black-on-white image that Tesseract processes with high accuracy.

Using Adobe Scan's page boundary controls before exporting

After Adobe Scan captures a page, it displays a crop view showing the detected page boundaries as adjustable handles. The automatic edge detection works well on white or near-white documents against a contrasting background, but struggles when the document color matches the surface it rests on. Manually drag each corner handle to align precisely with the document edge before accepting the capture. A crop that cuts into the text margin removes words that Tesseract would otherwise extract and redact. An overcropped edge is worse than an undercropped one for the redaction workflow: missing text cannot be redacted, while extra background in the image produces only minor OCR noise that does not interfere with the target field redaction.

Transferring Adobe Scan exports securely to a desktop for redaction

Adobe Scan exports to the device photo library when you disable cloud sync, but transferring that photo to a desktop machine for redaction requires choosing a transfer method that does not reintroduce the cloud exposure you disabled in the app. USB cable transfer via Finder on macOS or File Explorer on Windows copies the JPEG directly between devices over a wired connection with no cloud intermediary. AirDrop between Apple devices sends files directly over a peer-to-peer connection, making it a suitable local transfer method for sensitive documents.

Avoid sending the scan via email attachment or messaging app if those services route through cloud servers in jurisdictions where your document's sensitivity creates legal concerns. iMessage between Apple devices can use end-to-end encryption, but iCloud photo sync may capture the image before you send it if iCloud is enabled in the Files app. Google Photos auto-backup on Android similarly captures the scan before you can transfer it manually. Confirm that automatic backup is disabled in both the Adobe Scan settings and the device's photo management app before scanning any document that must remain local, then run it through the cloud-free Adobe Scan redaction flow.

Calibrating Adobe Scan for specific document types

Adobe Scan includes a Whiteboard capture mode and a Business Card mode alongside the default Document mode. The Document mode applies perspective correction and contrast enhancement suited for printed paper documents. For documents scanned inside clear plastic sleeves or binders, the plastic surface can create glare spots that appear as bright patches in the captured image, causing Tesseract to miss characters in the affected region. Remove the document from the sleeve before scanning, or photograph at a slight angle to push the glare to the document margin rather than the text region. For two-sided thermal paper receipts, which fade quickly, scan immediately after printing and use the Grayscale filter option under Adobe Scan color settings to preserve tonal variation that the standard Document mode might normalize away.

Sources
  1. 1.

    Adobe, "Adobe Scan," adobe.com, accessed June 2026. https://play.google.com/store/apps/details?id=com.adobe.scan.android&hl=en

  2. 2.

    Tesseract OCR, "Improving the quality of the output," tesseract-ocr.github.io, accessed June 2026. https://tesseract-ocr.github.io/tessdoc/ImproveQuality.html

  3. 3.

    Adobe, "Adobe Scan for Android," adobe.com, accessed June 2026. https://www.adobe.com/devnet-docs/adobescan/android/en/scan.html

  4. 4.

    Adobe, "Adobe Scan iOS Settings," adobe.com, accessed June 2026. https://www.adobe.com/devnet-docs/adobescan/ios/en/settings.html

  5. 5.

    ImageMagick, "Command-line Processing," imagemagick.org, accessed June 2026. https://imagemagick.org/command-line-processing/

FAQ

In the app, tap your profile avatar, go to Settings, and disable Auto-upload to Document Cloud. Also check whether your device camera roll syncs to iCloud or Google Photos, because Save to Camera Roll creates a second upload path through a different cloud service.

Adobe Scan's built-in OCR runs server-side and returns a text layer embedded in a PDF. The OCR Redactor runs Tesseract locally in the browser and exports a raster image with no text layer. The key difference is privacy: no file leaves your device during the process. For accuracy on clean scans, both engines produce similar results on standard printed text.

Yes. Export the scan as JPEG (Save Image), transfer it to any device with a browser, and open the OCR Redactor. The tool runs entirely offline after the initial page load; no file upload occurs. The full OCR and redaction workflow completes in the browser using the Tesseract WASM engine.

Adobe Scan applies automatic perspective correction and compression. For better downstream OCR, photograph under even, shadow-free lighting. Hold the camera directly overhead and parallel to the document surface to minimize perspective distortion. After export, CapyToolkit's OCR Redactor processes the JPEG regardless of original scan quality and can still extract most text from moderately blurry captures.

Adobe Scan still logs basic usage telemetry such as scan count and feature use. Disabling cloud sync stops the actual document content from uploading. For complete privacy, the OCR Redactor handles the entire OCR and redaction workflow locally with no document data leaving the browser at any point.

iPhone Document Scanner: OCR and Local Redaction Guide

iPhone's built-in scanner lives in the Files app and Notes. Open the Files app, tap the three-dot menu inside a folder, and choose Scan Documents to start a scan session in iOS 16 or later1. The scanner uses Apple's Vision framework on-device for perspective correction and tone adjustment, producing a JPEG without uploading anything to iCloud during the scan itself2.

JPEG quality defaults to a medium compression level in both apps. For the best OCR accuracy downstream, photograph with strong, even overhead lighting, minimize hand movement, and hold the phone at roughly 30 cm above the document3. Exports from the Files scanner default to JPEG; Notes exports to PDF, which needs conversion before the OCR Redactor can process it.

Run this check yourself in the Offline OCR & Document Redactor.

Open in the tool →

Specifications4

AvailabilityiOS 13+ (Files app, Notes, Control Center)
Output formatJPEG (Files app) or PDF (Notes)
OCR engineOn-device via Apple Vision framework
Export methodsAirDrop, Files share sheet, USB transfer
Cloud synciCloud Drive may sync if Files folder is in iCloud

Best capture settings for OCR accuracy

Photograph in natural or bright overhead light with no harsh shadows crossing text lines. The iPhone scanner uses Vision framework for real-time edge detection, but low-contrast scenes cause the edge detector to miss corners and crop incorrectly. Tap each corner handle in the crop view and drag it to align with the document edge before confirming.

Distance, resolution, and motion blur avoidance

Shooting at 30 cm distance from letter-size pages on a modern iPhone camera produces a result roughly equivalent to a 200 DPI scan3. For fonts below 10 points, move the phone to approximately 20 cm distance to increase effective resolution. Confirm no motion blur appears in the capture preview before accepting each page. Even slight hand movement at close range softens the image enough to degrade Tesseract accuracy on thin letterforms, so bracing your elbows against a table or using a small phone stand improves consistency across multi-page scans.

Common issues with iPhone scanner output

Notes exports documents as PDF rather than JPEG, which the OCR Redactor cannot accept directly. Use the Files app scanner instead, or take a plain screenshot of the Notes page and use that JPEG. Building on this, the iPhone camera applies tone mapping that boosts contrast but can over-sharpen fine text and introduce haloing around dark characters. For dense OCR work, disable Smart HDR under iOS Settings > Camera to get a more neutral tonal output. Furthermore, iCloud Drive sync may automatically upload the saved scan to Apple's servers if the Files app folder resides in iCloud Drive rather than On My iPhone5.

Feeding iPhone scans into the OCR Redactor

After scanning in the Files app, locate the saved JPEG in the designated local folder. Transfer it to a computer via AirDrop, a USB cable using the Photos import dialog, or by accessing a shared iCloud Drive folder from the computer. Drop the JPEG onto the OCR Redactor drop zone. Tesseract begins extracting text from the image. For scans with remaining perspective distortion, the extracted text may show spacing errors where characters were sampled at angles. Rescan with the phone held more directly overhead if spacing errors appear consistently. Yet most cleanly photographed documents under good lighting produce accurate OCR output on the first attempt.

Transferring iPhone document scans to a desktop without cloud services

Three transfer paths move an iPhone scan to a desktop without routing it through iCloud or a third-party cloud service. AirDrop works across macOS and iOS without any account: open the saved JPEG in Files, tap Share, select AirDrop, and choose the recipient Mac. On Windows, the Phone Link app built into Windows 11 receives photos from an iPhone via a Bluetooth or Wi-Fi direct connection, also without iCloud6.

Using a USB cable for direct file transfer

Connecting the iPhone to a Windows PC via Lightning or USB-C exposes the device as a camera in Windows Explorer under This PC. Open the device, browse to Internal Storage and then the DCIM folder, and copy the saved JPEG to the local file system. On macOS, Image Capture in Applications > Utilities imports photos directly to any folder you select. Both paths work with airplane mode enabled, so no network interface is required at transfer time.

A wired transfer keeps the document fully on your own hardware from capture to redaction. CapyToolkit's OCR Redactor reads the JPEG from the local file system and runs Tesseract in the browser, so the scan never contacts Apple servers or a third-party cloud during processing. For sensitive documents, complete the transfer with airplane mode enabled so no background sync can reach the file before you open the tool.

iOS Live Text vs. Tesseract browser OCR: privacy and accuracy comparison

iOS Live Text runs entirely on-device using Apple's Vision framework, extracting text from images in Photos without sending data to Apple servers. For quick lookups such as copying a phone number from a photo, Live Text is fast and requires no additional tools. For structured redaction with rectangles and a verified export, it has no equivalent workflow.

Why Tesseract in the browser adds redaction capability Live Text lacks

Tesseract.js running in the OCR Redactor also processes everything locally in your browser with no server upload. The key difference is the spatial bounding box interface: Tesseract returns word-level bounding boxes that the OCR Redactor uses to synchronize rectangle selections with text positions, the core of the Live Text versus Tesseract redaction comparison. When you draw a rectangle over a name in the image, the tool highlights the matching words in the text panel simultaneously. iOS Live Text does not expose this kind of synchronized highlight for redaction verification, making the OCR Redactor the more reliable audit tool when you need to confirm every instance of an identifier is covered before sharing a document.

Sources
  1. 1.

    Apple, "How to scan documents on your iPhone or iPad," support.apple.com, accessed June 2026. https://support.apple.com/en-us/108963

  2. 2.

    Apple, "Recognizing Text in Images," developer.apple.com, accessed June 2026. https://developer.apple.com/documentation/vision/recognizing-text-in-images

  3. 3.

    Genius Scan, "What is the DPI of my Scans?," help.geniusscan.com, accessed June 2026. https://help.geniusscan.com/other-topics/troubleshooting-and-faq/troubleshoot-scanning/what-is-the-dpi-of-my-scans

  4. 4.

    Apple, "Find files on your iPhone or iPad," support.apple.com, accessed June 2026. https://support.apple.com/en-us/102570

  5. 5.

    Microsoft, "Phone Link requirements and setup," support.microsoft.com, accessed June 2026. https://support.microsoft.com/en-us/topic/phone-link-requirements-and-setup-cd2a1ee7-75a7-66a6-9d4e-bf22e735f9e3

  6. 6.

    Tesseract.js, "API Documentation," github.com, accessed June 2026. https://github.com/naptha/tesseract.js/blob/master/docs/api.md

FAQ

The scanner appears in three places: Notes (camera icon > Scan Documents), Files (three-dot menu > Scan Documents in iOS 16 and later), and the iOS Control Center if you added it under Settings > Control Center. Files exports a JPEG while Notes exports a PDF. For the OCR Redactor, use the Files scanner to get a JPEG directly.

In the Files app, ensure your destination folder is On My iPhone rather than iCloud Drive. Tap the Browse tab, select On My iPhone, navigate to a local folder, and start the scan there. Files saved to On My iPhone folders do not sync to iCloud unless you share them separately.

Take a screenshot of the Notes scan as an alternative, or open the PDF in the Files app preview, tap Share, and choose Save Image to export the page as JPEG. The OCR Redactor accepts JPEG, PNG, and WebP images. For multi-page documents, save each page as a separate JPEG using the Files app preview.

Disable Smart HDR under Settings > Camera > Smart HDR. HDR tone mapping can add noise and haloing around small characters. Use the rear camera rather than the front camera for higher resolution. Photograph in bright, even lighting with no shadows across the text, and use the built-in scanner from the Files app for automatic perspective correction.

Yes. The OCR Redactor uses a WebAssembly build of Tesseract that works in Safari on iOS 15 and later. Tap the upload area and select your image from the Photos library. The full OCR and redaction workflow runs locally in the browser without uploading anything. CapyToolkit does not require an account or app installation.

FAQ

Use 300 DPI for normal printed documents. Move up to 400 or 600 DPI for small fonts, footnotes and dense tables, where letters are only a few pixels tall at 300 DPI. Going below 300 DPI saves space but usually costs accuracy on anything smaller than body text.

Usually, for single pages in good light. Phone scanning apps correct perspective and crop the page well, but their compression and the light in the room vary more than a document scanner's. Hold the phone steady and parallel to the page, and use a sheet-fed scanner for long or small-print documents.

As images, when the scan is going into the OCR Redactor. The redactor accepts JPEG, PNG and WebP, so a scanner set to PDF output means an extra export step for every page. Set the scanner to PNG or high-quality JPEG and keep the PDF assembly for after redaction.

Compression and resolution problems hide at normal zoom. Zoom the scan to 200 percent and look at thin letters such as i, l and t: blocky edges, smudged serifs or halos around the text mean the scanner's compression or resolution is too low for OCR, even if the page looks fine when you read it.

No. CapyToolkit doesn't connect to the scanner at all. You scan with your usual software or app, save the image to your device, and drop it into the OCR Redactor in your browser, where the text extraction and redaction run locally.