Skip to content

Evaluate browser PDF and OCR redaction #7

Description

@aanishs

Clinician value

Explore whether clinicians can redact simple PDFs in the browser without uploading documents.

Scope

  • Treat this as a second-phase evaluation after the text redactor is useful.
  • Prototype PDF.js text extraction and Tesseract.js OCR for image-based pages.
  • Document limitations: OCR misses text, redaction must be burned into output, and legal de-identification is not guaranteed.
  • Use synthetic documents only.

Acceptance checks

  • Evaluation notes state whether browser-only PDF redaction is reliable enough for a public tool.
  • No sample file contains real client information.
  • Prototype, if built, never uploads file contents.
  • The user-facing caveat is explicit before download/export.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions