Tips and tricks

Moving files to a “watched” directory

When moving an existing file to the directory FSCrawler is watching, you need to explicitly touch all the files as when moved, the files are keeping their original date intact:

# single file
touch file_you_moved

# all files
find  -type f  -exec touch {} +

# all .txt files
find  -type f  -name "*.txt" -exec touch {} +

Or you need to restart from the beginning with the --restart option which will reindex everything.

Workaround for huge temporary files

FSCrawler uses a media library that currently does not clean up their temporary files. Parsing MP4 files may create very large temporary files in /tmp. The following commands could be useful e.g. as a cronjob to automatically delete those files once they are old and no longer in use. Adapt the commands as needed.

# Check all files in /tmp
find /tmp \( -name 'apache-tika-*.tmp-*' -o -name 'MediaDataBox*' \) -type f -mmin +15 ! -exec fuser -s {} \; -delete

# When using a systemd service with PrivateTMP enabled
find $(find /tmp -maxdepth 1 -type d -name 'systemd-private-*-fscrawler.service-*') \( -name 'apache-tika-*.tmp-*' -o -name 'MediaDataBox*' \) -type f -mmin +15 ! -exec fuser -s {} \; -delete

Indexing from HDFS drive

There is no specific support for HDFS in FSCrawler. But you can mount your HDFS on your machine and run FS crawler on this mount point. You can also read details about HDFS NFS Gateway.

Using docker

See Using docker.

Running FSCrawler on multiple machines

Added in version 3.0.

If you run FSCrawler on several hosts against the same Elasticsearch cluster, and those hosts crawl paths that look identical (for example both watch /data/docs), document _ids can collide.

By default, the _id is derived from the file path (see Document IDs). The same path on machine1 and machine2 therefore produces the same _id. The last writer wins and silently overwrites the other machine’s document.

Do not point several crawlers at the same physical document index unless every path (and thus every _id) is guaranteed unique across machines.

Alternatives

  • Different fs.url layouts that never collide in the hashed path (for example mount points that include the hostname) can share one index, but that is brittle and hard to reason about.

  • fs.filename_as_id: true does not fix multi-machine collisions: identical filenames still share the same _id.

See also

Deduplicate documents with a content-based _id

By default, FSCrawler derives the Elasticsearch document _id from the file path (see Document IDs). Two identical files under different paths therefore become two distinct documents.

If you want to index only one copy of duplicate files, you can set the document _id from a content fingerprint via an Elasticsearch ingest pipeline.

Option 1: binary checksum (fs.checksum)

Enable file checksum so FSCrawler stores a hash of the binary file in file.checksum:

name: "test"
fs:
  index_content: true
  # indexed_chars: 0   # optional: checksum only, no extracted text
  checksum: "SHA-256"
elasticsearch:
  pipeline: "set-id-from-checksum"

Create the pipeline that copies file.checksum into _id:

PUT _ingest/pipeline/set-id-from-checksum
{
  "description": "Set the _id from file.checksum",
  "processors": [
    {
      "set": {
        "field": "_id",
        "value": "{{{file.checksum}}}"
      }
    }
  ]
}

Identical binary files then share the same _id. With the default elasticsearch.bulk_operation: index, the last indexed copy wins and overwrites the previous document. That behaviour is useful when you update a file and want the index to reflect the latest path or metadata for that content.

To keep the first indexed copy instead, set elasticsearch.bulk_operation: create:

name: "test"
fs:
  index_content: true
  checksum: "SHA-256"
elasticsearch:
  pipeline: "set-id-from-checksum"
  bulk_operation: "create"

Later duplicates fail with a create conflict that FSCrawler treats as expected (non-fatal), so a single document remains and its version stays at 1.

Note

The checksum is computed from the binary file, not from the extracted text. fs.checksum is independent from fs.hash_algorithm (the latter only affects path-based ids).

Option 2: fingerprint of the extracted text

If you prefer to deduplicate on the extracted text (content) instead of the binary, use the Elasticsearch fingerprint ingest processor:

PUT _ingest/pipeline/content-fingerprint-id
{
  "description": "Compute a fingerprint from content and set it as the document _id",
  "processors": [
    {
      "fingerprint": {
        "fields": ["content"],
        "target_field": "_tmp_fingerprint",
        "method": "SHA-256"
      }
    },
    {
      "set": {
        "field": "_id",
        "value": "{{{_tmp_fingerprint}}}"
      }
    },
    {
      "remove": {
        "field": "_tmp_fingerprint",
        "ignore_missing": true
      }
    }
  ]
}

Then set elasticsearch.pipeline: "content-fingerprint-id" in your job settings.

Warning

This option overwrites documents that have no extracted text or the exact same text. Binary-identical files with different extracted text are not treated as duplicates. Conversely, different files that yield the same extracted text are treated as duplicates.

Caveats

  • bulk_operation: with the default index, when several paths share the same fingerprint only one document remains and the last writer wins. With create, the first writer wins and later duplicates are skipped (see Elasticsearch settings).

  • fs.remove_deleted: deleting one of the duplicate files on disk can remove the shared document from Elasticsearch even if another copy still exists.

  • Folder documents are not sent through the ingest pipeline (see Using Ingest Node Pipeline).

  • Changing to a content-based _id produces new ids; reindex from a clean index (or with --restart) as described in Document IDs.

See also