Kreuzberg
This module extracts text and Markdown from uploaded documents, supporting various formats and OCR for scanned files. It provides cached results with page information and metadata, designed to be used by other modules for tasks like indexing or processing document content with AI.
Kreuzberg extracts text and Markdown from documents attached to Drupal file entities, using an Xberg server (formerly Kreuzberg). Results keep page boundaries, document metadata and an OCR flag, and are cached per file content.
It is an API module: other modules call the extractor service, for example to hand a contract or an invoice to an LLM, or to index document text.
Features
- PDF, DOCX and the other formats Xberg supports, with OCR for scanned documents
- Markdown or plain-text output, per page and for the whole document
- Works with any stream wrapper, including
private:// - Results are cached by file UUID, file content hash and options, and invalidated when the file changes
- Status report entry showing whether the server is reachable and which Xberg version it runs
Requirements
- Drupal 10.3 or 11, PHP 8.3+
- A running Xberg HTTP server. The official image is
ghcr.io/xberg-io/xberg; 1.0.x is tested against 1.3.2.
Running Xberg
Run Xberg as a sidecar container on an internal network only: it accepts any file and has no authentication. Pin a version tag rather than latest.
services:
xberg:
image: ghcr.io/xberg-io/xberg:1.3.2
restart: unless-stopped
expose:
- "8000"
volumes:
- xberg-cache:/app/.xberg
volumes:
xberg-cache:
For DDEV, put the service into .ddev/docker-compose.xberg.yaml.
Configuration
Set the server URL (default http://xberg:8000) and timeout at Configuration > Media > Kreuzberg. This needs the Administer Kreuzberg permission. You can also set the URL per environment in settings.php:
$config['kreuzberg.settings']['server_url'] = getenv('XBERG_URL');
Usage
$result = \Drupal::service('kreuzberg.extractor')->extract($file);
$result->content; // Markdown of the whole document.
$result->pages; // [['number' => 1, 'content' => '...'], ...]
$result->pageCount; // NULL for formats without pages.
Failures throw a subclass of KreuzbergException:
ExtractionUnavailableException: the server is down or timed outUnsupportedFormatException: Xberg does not support the formatEmptyExtractionException: extraction found no text
The extractor does not check file access; that is the caller's job. Keep your own allowlist of accepted formats: Xberg reads many formats, and a high quality score does not guarantee useful text (Apple Pages files, for example, yield only a preview).
Roadmap
- An in-process backend based on the Xberg PHP extension, as an alternative to the HTTP server
Developed and maintained by Factorial.
Depends on
Dependencies of the latest stable release
No dependencies recorded for this project.
Required by
Tracked projects that depend on this one
No tracked projects depend on this one yet.
Activity
Releases
| Version | Type | Core | Notes | Release date | |
|---|---|---|---|---|---|
| 1.0.x-dev | Dev | 10–11 | Oct 2, 2026 |