<a id="release-notes-3-0"></a>

# Version 3.0

#### IMPORTANT
There is no in-place upgrade from FSCrawler 2.9. Install 3.0, recreate jobs with
`--setup`, and reindex. See [Upgrading from 2.9](../user/upgrade.html.md#upgrade-from-2-9).

## Breaking changes

- If you want to exclude a specific folder, you need to use a wildcard character at the end of the folder name.
  For example, to exclude the folder `/tmp/foo`, you need to use `/tmp/foo/*`. Thanks to dadoonet.
- The way we run docker images has changed. We don’t need anymore to specify the fscrawler binary.
  So running `docker run -it -v ~/.fscrawler:/root/.fscrawler -v /documents:/tmp/es:ro dadoonet/fscrawler job_name` is
  enough. Thanks to dadoonet.
- FSCrawler does not display anymore the list of existing jobs when no job name is provided.
  You need to use the `--list` option to list the jobs. Thanks to dadoonet.
- When launching for the first time FSCrawler with a job name, FSCrawler does not create anymore the job
  configuration folder with default settings. You need to use the `--setup` option to create the job settings.
  Thanks to dadoonet.
- We don’t support anymore the `elasticsearch.nodes.url` setting. You need to use `elasticsearch.urls`
  instead. Thanks to dadoonet.
- The `_upload` REST endpoint has been removed. Please now use the `_document` endpoint. Thanks to dadoonet.
- The Apache Tika 4 upgrade changes several `meta.raw.*` metadata key names compared to FSCrawler 2.9.
  If you search, aggregate or map these fields, you will need to update your queries and index templates
  accordingly:
  - Image and EXIF metadata (JPEG, PNG, TIFF, …) keys are now namespaced under `img:`. For example
    `Number of Tables` becomes `img:Number of Tables` and `Exif IFD0:Orientation` becomes
    `img:Exif IFD0:Orientation`.
  - ICC profile keys use the lowercase `icc:` prefix instead of `ICC:`.
  - PDF `access-permission:*` and `pdf:*` keys now use hyphens instead of underscores or camelCase.
    For example `access_permission:assemble_document` becomes `access-permission:assemble-document`
    and `pdf:PDFVersion` becomes `pdf:pdf-version`.
  - The resource name key `resourceName` is renamed to `tk:resource-name`.
  - Tika internal keys now use the `tk:` prefix instead of `X-TIKA:`. New keys include
    `tk:content-type-magic-detected`, `tk:parsed-by-full-set` and, for text documents, the encoding
    detection keys `tk:detected-encoding`, `tk:encoding-detection-trace` and `tk:encoding-detector`.
    Thanks to dadoonet.
- The default `fs.ocr.pdf_strategy` is now `auto` instead of `ocr_and_text`. With `auto`, OCR is skipped on
  PDF pages that already contain more than 10 characters of text. If you relied on the previous behaviour and want
  OCR on every PDF page, explicitly set `fs.ocr.pdf_strategy` to `ocr_and_text`. Thanks to dadoonet.
- New jobs created with `--setup` set `fs.hash_algorithm` to `SHA-256` in the example settings. Existing jobs that
  omit the setting keep `MD5` so document `_id`s stay unchanged. Changing the algorithm later requires a full
  reindex. See [Document IDs](../admin/fs/document-ids.html.md#document-ids). Closes [#2425](https://github.com/dadoonet/fscrawler/issues/2425). Thanks to dadoonet.

## New

- Default `fs.excludes` now also skips macOS Finder metadata files (`.DS_Store`), via the
  case-insensitive pattern `*/.ds_store`, in addition to `*/~*`. See [Includes and excludes](../admin/fs/local-fs.html.md#includes-excludes).
  Thanks to dadoonet.
- The crawler system has been unified using a plugin architecture. You can now specify the crawler provider using
  `fs.provider` instead of `server.protocol`. Available providers are `local` (default), `ftp`, and `ssh`.
  See [Crawler Provider](../admin/fs/local-fs.html.md#crawler-provider). Thanks to dadoonet.
- FSCrawler does not need to wait until the next planned scan to scan again the filesystem. You can just set the
  `next_check` field to `null` in the `~/.fscrawler/{job_name}/_checkpoint.json` file and FSCrawler will start
  a new scan immediately.
- Job settings can be defined by env variables and system properties and you can also split the configuration of
  jobs using multiple files in the `~/.fscrawler/job/_settings` directory. Also note that the system properties
  need to be set in the `FS_JAVA_OPTS` environment variable.
- Add support for automatic semantic search when using a 8.17+ version with a trial or enterprise
  license. See [Semantic search](../admin/fs/elasticsearch.html.md#semantic-search). **Warning**: this might slow down the ingestion process. Thanks to dadoonet.
- Add support for Elastic cloud serverless. Thanks to dadoonet.
- Using the REST API `_document`, you can now fetch a document from the local dir, from an http website
  or from an S3 bucket. See [REST service](../admin/fs/rest.html.md#rest-service). Thanks to dadoonet.
- You can now remove a document in Elasticsearch using FSCrawler `_document` endpoint. See [REST service](../admin/fs/rest.html.md#rest-service). Thanks to dadoonet.
- Implement our own HTTP Client for Elasticsearch. Thanks to dadoonet.
- FSCrawler now ships with Apache Tika’s `tika-vlm` module: OCR can be delegated to a Vision Language
  Model through an OpenAI-compatible endpoint (vLLM, Ollama, Azure OpenAI…), Anthropic Claude or Google
  Gemini, configured via a custom Tika configuration file. See [Using a Vision Language Model (VLM) for OCR](../user/ocr.html.md#vlm-ocr). The default FSCrawler
  parser chain is unchanged (Tesseract when available). Thanks to dadoonet.
- Add option to set path to custom tika config file. See [Local FS settings](../admin/fs/local-fs.html.md#local-fs-settings). Thanks to iadcode for the original
  XML implementation and to betofilippi for the switch to JSON.
  If your JSON configuration uses `default-parser`, exclude the VLM parser components unless you
  explicitly enable one — see [Local FS settings](../admin/fs/local-fs.html.md#local-fs-settings) and [Using a Vision Language Model (VLM) for OCR](../user/ocr.html.md#vlm-ocr).
  **Note**: since the Apache Tika 4 upgrade, the configuration file must be a Tika JSON configuration —
  the XML-based configuration file mechanism was removed upstream. Existing XML configurations need to be
  converted. See [Local FS settings](../admin/fs/local-fs.html.md#local-fs-settings).
- Support for Index Templates. See [Mappings](../admin/fs/elasticsearch.html.md#mappings). Thanks to dadoonet.
- Support for Aliases. You can now index to an alias. Thanks to dadoonet.
- Support for Access Token and Api Keys instead of Basic Authentication. See [Using Credentials (Security)](../admin/fs/elasticsearch.html.md#credentials). Thanks to dadoonet.
- Allow loading external jars. This adds a new `external` directory from where jars can be loaded
  to the FSCrawler JVM. For example, you could provide your own Custom Tika Parser code. See [Directory layout](../admin/layout.html.md#layout). Thanks to dadoonet.
- Add temporal information in folder index. Thanks to bdauvissat
- Add support for external metadata files while crawling, defaults to `.meta.yml`. See [External Tags](../admin/fs/tags.html.md#tags) Thanks to dadoonet.
- Add support for static external metadata for all documents. See [External Tags](../admin/fs/tags.html.md#tags) Thanks to dadoonet.
- The job name is not mandatory anymore and it will be `fscrawler` by default. Thanks to dadoonet.
- FSCrawler also supports Elasticsearch 9. Thanks to dadoonet.
- Add support for ACL metadata extraction for NTFS filesystems, including principals, permissions, and flags. Thanks to alexbluesteele.
- Add support for pause/resume functionality with checkpoint persistence. The crawler can now be paused and resumed
  without losing progress. It also automatically recovers from network errors with exponential backoff retry.
  See [REST service](../admin/fs/rest.html.md#rest-service). Thanks to dadoonet.
- HTTP retry backoff is configurable via `elasticsearch.retry_max_duration`,
  `elasticsearch.retry_initial_delay` and `elasticsearch.retry_max_delay`
  (defaults `5m` / `500ms` / `30s`). The same budget applies to `5xx`, `429`, and a
  cold-start `404` on `GET /`. See [HTTP retry settings](../admin/fs/elasticsearch.html.md#http-retry-settings). Thanks to dadoonet.
- FSCrawler can create a default Kibana dashboard on job startup via the Kibana Dashboards API (Kibana 9.5+).
  See [Kibana settings](../admin/fs/kibana.html.md#kibana-settings). Closes [#2477](https://github.com/dadoonet/fscrawler/issues/2477). Thanks to dadoonet.

## Fix

- Apple Keynote (`.key`) files are now supported for content extraction and indexing. Closes [#782](https://github.com/dadoonet/fscrawler/issues/782).
- Closed open file streams after use. Thanks to alexbluesteele.
- `fs.ocr.enabled` was always false. Thanks to ywjung.
- Do not hide YAML parsing errors. Thanks to dadoonet.
- Fix duration parsing for the day unit `d`. Thanks to dadoonet.
- Image raw metadata extraction was not working. Thanks to dadoonet.
- Fix issue when using crawling over SSH when the directory ends with a space. Thanks to dadoonet.
- On Windows, files and directories to be removed were not properly detected. Thanks to newschapmj1.
- Bulk `_bulk` HTTP calls now retry on `429`/`5xx` and no longer treat a failed bulk as success.
  Exhausted retries mark the crawl checkpoint as `ERROR` and REST uploads return `ok: false`.
  Thanks to dadoonet.
- Default Log4J config sets `org.apache.pdfbox` to `error` to avoid flooding logs with
  `No Unicode mapping for CID+…` warnings from subset fonts. Thanks to dadoonet.
- Default Log4J config sets `org.apache.fontbox` to `error` to avoid flooding logs with
  `No PostScript name data is provided for the font …` warnings. Thanks to dadoonet.
- Failed bulk actions are detailed in `logs/bulk-failures.log` (reason-prefixed lines; truncated
  payloads at TRACE). Console / `fscrawler.log` point to that file. Thanks to dadoonet.

## Deprecated

- The `server.protocol` setting is deprecated. Use `fs.provider` instead. Thanks to dadoonet.
- Support for Basic Authentication is deprecated. You should use API keys instead. Thanks to dadoonet.

## Updated

- Files are now sorted by date with a reverse order. So the most recent files should be indexed first. Thanks to dadoonet.
- Add full support for Elasticsearch [9.5.2](https://www.elastic.co/docs/solutions/search), [8.19.5](https://www.elastic.co/guide/en/elasticsearch/reference/8.19/index.html), [7.17.29](https://www.elastic.co/guide/en/elasticsearch/reference/7.17/index.html). Thanks to dadoonet.
- Update to [Tika 4.0.0](https://tika.apache.org/4.0.0/). Thanks to dadoonet.
- The default alias name is now the job name and not forced to `fscrawler` anymore. Thanks to dadoonet.
- The default REST endpoint is now running at `/` instead of `/fscrawler/`. Thanks to dadoonet.
- Upgrade to Jackson 3.x. Closes [#2419](https://github.com/dadoonet/fscrawler/issues/2419). Thanks to dadoonet.
- As a consequence of the Jackson 3 upgrade, the JSON and YAML documents produced by FSCrawler now serialize their
  fields in alphabetical order. This affects the indexed documents, the generated `_settings.yaml` and
  `_checkpoint.json` files, and the REST API responses. This is purely cosmetic (field order is not significant in
  JSON), but users who version their configuration files may notice a one-time reordering. Thanks to dadoonet.

## Removed

- Remove the specific distributions depending on Elastic version. Thanks to dadoonet.
- Support for Elasticsearch 6.x is removed. Thanks to dadoonet.

Thanks to `@dadoonet`, `@ywjung`, `@iadcode`, `@bdauvissat`, `@alexbluesteele`, `@betofilippi`, `@newschapmj1`
and all the contributors for this release!
