Skip to content

[Improvement] Add configuration to stop parsing PDFs after X pages #1901

Description

@jnioche

What would you like to be improved?

Every so often an open crawl will stumble upon a very large pdf, these can take a lot of CPU to parse when effectively most of the content will not be indexed.

How should we improve?

apache/tika#2803 introduced a config for PDF parsing in Tika to stop processing after X pages. We should make use of it as soon as the next version of Tika is released (currently 3.3.0)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions