Skip to content

C API Reference

Extract content from a single bytes or URI input.

Signature:

XBERGAlefHandle xberg_extract(XBERGAlefHandle input, XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_extract(0, 0);

Parameters:

Name Type Required Description
input XBERGAlefHandle Yes The input data
config XBERGAlefHandle Yes The configuration options

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.


Extract content from multiple bytes or URI inputs.

Signature:

XBERGAlefHandle xberg_extract_batch(const char* inputs, XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_extract_batch(NULL, 0);

Parameters:

Name Type Required Description
inputs const char* Yes The inputs
config XBERGAlefHandle Yes The configuration options

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.


Discover all pages and sitemaps reachable from uri without extracting document content.

Builds a crawlberg.CrawlEngine from config.crawl, calls CrawlEngine.map, and returns the set of discovered URLs as a crawlberg.MapResult (re-exported as MapResult).

Use this when you need the URL inventory of a site before committing to full document extraction — e.g. to build a crawl queue or validate scope.

Errors:

Returns Validation if the crawl configuration fails validation or if the map operation itself fails.

Signature:

XBERGAlefHandle xberg_map_url(const char* uri, XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_map_url("value", 0);

Parameters:

Name Type Required Description
uri const char* Yes The uri
config XBERGAlefHandle Yes The configuration options

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.


List all supported document formats.

Returns every file extension Xberg recognizes together with its corresponding MIME type, derived from the central format registry. Formats that have no registered file extension (such as source code, which is detected dynamically) are not included.

The static EXT_TO_MIME table lists every format the codebase knows how to describe, regardless of which Cargo features were compiled in. Advertising that table directly would claim support for extractors that may not exist in this build (see GH#1387). To keep the advertised catalogue honest, the table is intersected with the document extractor registry: an extension is only included if some registered extractor actually claims its MIME type in this build. This can never drift from reality and automatically covers third-party extractors registered at runtime.

The list is sorted alphabetically by file extension.

Returns:

A vector of SupportedFormat entries sorted by extension, limited to formats with a registered extractor in this build.

Signature:

const char* xberg_list_supported_formats();

Example:

const char* result = xberg_list_supported_formats();

Returns: const char*


Clear all embedding backends from the global registry.

Calls shutdown() on every registered backend, then empties the registry.

Errors:

  • Any error returned by a backend’s shutdown() method. The first error encountered stops processing of remaining backends.

Signature:

int32_t xberg_clear_embedding_backends();

Example:

xberg_clear_embedding_backends();

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


List the names of all registered embedding backends.

Used by xberg-cli, the api/mcp endpoints, and generated language bindings.

Signature:

const char* xberg_list_embedding_backends();

Example:

const char* result = xberg_list_embedding_backends();

Returns: const char*

Errors: Returns NULL on error.


List names of all registered document extractors.

Signature:

const char* xberg_list_document_extractors();

Example:

const char* result = xberg_list_document_extractors();

Returns: const char*

Errors: Returns NULL on error.


Clear all document extractors from the global registry.

Calls shutdown() on every registered extractor, then empties the registry.

Errors:

  • Any error returned by an extractor’s shutdown() method. The first error encountered stops processing of remaining extractors.

Signature:

int32_t xberg_clear_document_extractors();

Example:

xberg_clear_document_extractors();

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


List all registered OCR backends.

Returns the names of all OCR backends currently registered in the global registry.

Returns:

A vector of OCR backend names.

Signature:

const char* xberg_list_ocr_backends();

Example:

const char* result = xberg_list_ocr_backends();

Returns: const char*

Errors: Returns NULL on error.


List every registered OCR backend’s name alongside its declared supported languages.

This is the capability-enumeration counterpart to list_ocr_backends: where that function exposes only backend names, this exposes each backend’s supported_languages() too, so a consumer (for example, a job-acceptance gate) does not need to hardcode a second list of backend languages.

The returned vector is sorted by name, regardless of registration order or the order reported by the underlying registry. supported_languages within each entry is not sorted — see OcrBackendCapabilities.supported_languages for why.

Calling this is not free for every backend. In particular, TesseractBackend’s supported_languages() allocates a Tesseract API and initializes it against the same tessdata directory a real OCR job resolves (resolve_tessdata_path), the first time it is called, to enumerate installed tessdata languages; subsequent calls are served from a cache.

Errors:

Returns an error only if the registry lock cannot be acquired in the current environment.

Signature:

const char* xberg_list_ocr_backend_capabilities();

Example:

const char* result = xberg_list_ocr_backend_capabilities();

Returns: const char*

Errors: Returns NULL on error.


list_ocr_backend_capabilities, reporting each backend’s languages under config.

Tesseract’s language list is a property of the tessdata directory it resolves, and config.tessdata_path is the first entry of that search chain. The config-less form always answers for the no-override chain, which can be a different directory than the one a job using config will load from. Use this form when config.tessdata_path is set. See GH#1857.

Signature:

const char* xberg_list_ocr_backend_capabilities_for(XBERGAlefHandle config);

Example:

const char* result = xberg_list_ocr_backend_capabilities_for(0);

Parameters:

Name Type Required Description
config XBERGAlefHandle Yes The configuration options

Returns: const char*

Errors: Returns NULL on error.


Check whether a specific registered OCR backend supports a language.

Delegates to the named backend’s own OcrBackend.supports_language, which is the correct per-language decision — do not infer support (or its absence) from whether list_ocr_backend_capabilities reports an empty supported_languages list for that backend, since an empty list can mean “does not enumerate” rather than “supports nothing” (see OcrBackendCapabilities.supported_languages).

Errors:

Returns an error if no backend with that name (or alias) is registered.

Signature:

int32_t xberg_ocr_backend_supports_language(const char* backend, const char* language);

Example:

int32_t result = xberg_ocr_backend_supports_language("value", "value");

Parameters:

Name Type Required Description
backend const char* Yes Name of a registered OCR backend, as returned by list_ocr_backends. Lookup is case-insensitive and resolves the same paddleocr alias as backend dispatch.
language const char* Yes Language code to check (e.g. "eng", "deu").

Returns: int32_t

Errors: Returns 0 on error.


ocr_backend_supports_language, answering under config.

Use this, not the config-less form, when the caller sets OcrConfig.tessdata_path: the config-less form checks the no-override search chain and can deny a language the job would load without trouble. See GH#1857.

Signature:

int32_t xberg_ocr_backend_supports_language_for(const char* backend, const char* language, XBERGAlefHandle config);

Example:

int32_t result = xberg_ocr_backend_supports_language_for("value", "value", 0);

Parameters:

Name Type Required Description
backend const char* Yes The backend
language const char* Yes The language
config XBERGAlefHandle Yes The configuration options

Returns: int32_t

Errors: Returns 0 on error.


Clear all OCR backends from the global registry.

Removes all OCR backends and calls their shutdown() methods.

Returns:

  • Ok(()) if all backends were cleared successfully
  • Err(...) if any shutdown method failed

Signature:

int32_t xberg_clear_ocr_backends();

Example:

xberg_clear_ocr_backends();

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


List all registered post-processor names.

Returns a vector of all post-processor names currently registered in the global registry.

Returns:

  • Ok(const char**) - Vector of post-processor names
  • Err(...) if the registry lock is poisoned

Signature:

const char* xberg_list_post_processors();

Example:

const char* result = xberg_list_post_processors();

Returns: const char*

Errors: Returns NULL on error.


Remove all registered post-processors.

The next post-processed extraction restores enabled built-in processors before it snapshots the processor cache. Custom processors remain removed. Use unregister_post_processor when one named processor should remain absent while the rest of the registry stays intact. Returns a retryable in-use error when an extraction is executing a processor snapshot.

Signature:

int32_t xberg_clear_post_processors();

Example:

xberg_clear_post_processors();

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


List names of all registered renderers.

Errors:

Returns an error if the registry lock is poisoned.

Signature:

const char* xberg_list_renderers();

Example:

const char* result = xberg_list_renderers();

Returns: const char*

Errors: Returns NULL on error.


Clear all renderers from the global registry.

Removes every renderer, including the built-in defaults (markdown, html, djot, plain). After calling this no renderers are registered; re-register as needed.

Errors:

Returns an error if the registry lock is poisoned.

Signature:

int32_t xberg_clear_renderers();

Example:

xberg_clear_renderers();

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


Clear all reranker backends from the global registry.

Calls shutdown() on every registered backend, then empties the registry.

Errors:

  • Any error returned by a backend’s shutdown() method. The first error encountered stops processing of remaining backends.

Signature:

int32_t xberg_clear_reranker_backends();

Example:

xberg_clear_reranker_backends();

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


List the names of all registered reranker backends.

Used by xberg-cli, the api/mcp endpoints, and generated language bindings.

Signature:

const char* xberg_list_reranker_backends();

Example:

const char* result = xberg_list_reranker_backends();

Returns: const char*

Errors: Returns NULL on error.


Clear all tokenizer backends from the global registry.

Calls shutdown() on every registered backend, then empties the registry.

Errors:

  • Any error returned by a backend’s shutdown() method. The first error encountered stops processing of remaining backends.

Signature:

int32_t xberg_clear_tokenizer_backends();

Example:

xberg_clear_tokenizer_backends();

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


List the names of all registered tokenizer backends.

Used by xberg-cli, the api/mcp endpoints, and generated language bindings.

Signature:

const char* xberg_list_tokenizer_backends();

Example:

const char* result = xberg_list_tokenizer_backends();

Returns: const char*

Errors: Returns NULL on error.


List names of all registered validators.

Signature:

const char* xberg_list_validators();

Example:

const char* result = xberg_list_validators();

Returns: const char*

Errors: Returns NULL on error.


Remove all registered validators.

Signature:

int32_t xberg_clear_validators();

Example:

xberg_clear_validators();

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


Classify a document’s chunks and return the updated document.

This owned form preserves the mutations when the document crosses a language-binding boundary. Rust callers that already own a mutable document can use classify_chunks to avoid moving it.

Errors:

Returns the same validation and LLM errors as classify_chunks.

Signature:

XBERGAlefHandle xberg_classify_chunks_owned(XBERGAlefHandle result, XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_classify_chunks_owned(0, 0);

Parameters:

Name Type Required Description
result XBERGAlefHandle Yes The extracted document
config XBERGAlefHandle Yes The configuration options

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.


Find unmarked claims in markdown text.

Returns lines that assert a claim but carry neither a footnote citation anchor ([^...]) nor an inference marker ([*inference*]).

The heuristic is simple: a line that contains alphabetic words, ends with sentence punctuation, and is not a heading, blank line, or markup-only line is considered a claim. Exclude lines that appear in the citation block (after --- + <!-- citations ... -->).

Returns:

A vector of trimmed line text strings for unmarked claims.

Signature:

const char* xberg_find_unmarked_claims(const char* markdown);

Example:

const char* result = xberg_find_unmarked_claims("value");

Parameters:

Name Type Required Description
markdown const char* Yes The markdown text to search

Returns: const char*


Verify that an excerpt appears verbatim in source text.

Performs exact matching by default. Also tries whitespace-normalized matching (collapsing runs of whitespace on both sides) since PDF-extracted text often has irregular spacing.

Returns:

true if the excerpt appears (exactly or with normalized whitespace), false otherwise.

Signature:

int32_t xberg_verify_excerpt(const char* excerpt, const char* source_text);

Example:

int32_t result = xberg_verify_excerpt("value", "value");

Parameters:

Name Type Required Description
excerpt const char* Yes The text snippet to find
source_text const char* Yes The full source text to search

Returns: int32_t


Score a query against a document using ColBERT’s MaxSim operator: for each query token vector, take the maximum dot product against any document token vector, then sum across query tokens.

Returns 0.0 if query and doc have mismatched dimensionality, if either has zero tokens, or if either is not well-formed per MultiVectorEmbedding.is_well_formed (its data length does not match num_tokens * dim).

Pure CPU primitive — available without ONNX Runtime.

Signature:

double xberg_max_sim_score(XBERGAlefHandle query, XBERGAlefHandle doc);

Example:

double result = xberg_max_sim_score(0, 0);

Parameters:

Name Type Required Description
query XBERGAlefHandle Yes The multi vector embedding
doc XBERGAlefHandle Yes The multi vector embedding

Returns: double


Rank a set of documents against a query by MaxSim score, descending.

Mirrors the sort/truncate shape of crate.reranking’s build_results, minus top-k truncation (callers slice the returned Vec themselves).

Pure CPU primitive — available without ONNX Runtime.

Signature:

const char* xberg_max_sim_rank(XBERGAlefHandle query, const char* docs);

Example:

const char* result = xberg_max_sim_rank(0, NULL);

Parameters:

Name Type Required Description
query XBERGAlefHandle Yes The multi vector embedding
docs const char* Yes The docs

Returns: const char*


Probe the backends and settings in config and report what will actually execute on this host.

Runs no downloads and no billable API calls. Backends that are not compiled in or whose models are not cached report Skip rather than failing.

Signature:

XBERGAlefHandle xberg_doctor(XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_doctor(0);

Parameters:

Name Type Required Description
config XBERGAlefHandle Yes The configuration options

Returns: XBERGAlefHandle


Count the pages in a PDF without rendering any of them.

Opens the document and returns its page count from the PDF structure. No page is rasterized, so this is cheap relative to render_pdf_page_to_png — use it when you only need the count (e.g. to drive a render loop over the pages).

Errors:

Returns XbergError.Parsing if the PDF cannot be opened, authenticated, or its page count read.

Signature:

uintptr_t xberg_pdf_page_count(const uint8_t* pdf_bytes, const char* password);

Example:

uintptr_t result = xberg_pdf_page_count((const uint8_t *)"data", "value");

Parameters:

Name Type Required Description
pdf_bytes const uint8_t* Yes Raw PDF file bytes
password const char* No Optional password for encrypted PDFs

Returns: uintptr_t

Errors: Returns 0 on error.


C representation: XBERGAccelerationConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGAccelerationConfig does not appear anywhere in the generated header.

Hardware acceleration configuration for ONNX Runtime models.

Controls which execution provider (CPU, CoreML, CUDA, TensorRT) is used for inference in layout detection and embedding generation.

Field Type Default Description
provider XBERGAlefHandle XBERG_AUTO Execution provider to use for ONNX inference.
device_id uint32_t — GPU device ID (for CUDA/TensorRT). Ignored for CPU/CoreML/Auto.

C representation: XBERGArchiveEntry is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGArchiveEntry does not appear anywhere in the generated header.

A single file extracted from an archive.

When archives (ZIP, TAR, 7Z, GZIP) are extracted with recursive extraction enabled, each processable file produces its own full ExtractedDocument.

Field Type Default Description
path const char* — Archive-relative file path (e.g. “folder/document.pdf”).
mime_type const char* — Detected MIME type of the file.
result XBERGAlefHandle — Full extraction result for this file.

C representation: XBERGArchiveMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGArchiveMetadata does not appear anywhere in the generated header.

Archive (ZIP/TAR/7Z) metadata.

Extracted from compressed archive files containing file lists and size information.

Field Type Default Description
format const char* — Archive format (“ZIP”, “TAR”, “7Z”, etc.)
file_count uint32_t — Total number of files in the archive
file_list const char* NULL List of file paths within the archive
total_size uint64_t — Total uncompressed size in bytes
compressed_size uint64_t* NULL Compressed size in bytes (if available)

C representation: XBERGAttributes is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGAttributes does not appear anywhere in the generated header.

Element attributes in Djot.

Represents the attributes attached to elements using {.class #id key=“value”} syntax.

Field Type Default Description
id const char* NULL Element ID (#identifier)
classes const char* NULL CSS classes (.class1 .class2)
key_values const char* NULL Key-value pairs (key=“value”)

C representation: XBERGAudioMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGAudioMetadata does not appear anywhere in the generated header.

Audio/video file metadata.

Populated from container tags (ID3v2, MP4 atoms, Vorbis comments, etc.) and PCM decode properties. Available when the transcription-types feature is enabled.

Field Type Default Description
duration_ms uint64_t* NULL Duration in milliseconds derived from the decoded audio stream.
codec const char* NULL Audio codec (e.g. “mp3”, “aac”, “opus”, “flac”).
container const char* NULL Container format (e.g. “mpeg”, “mp4”, “ogg”, “wav”).
sample_rate_hz uint32_t* NULL Sample rate in Hz after decode (always 16000 when resampled for Whisper).
channels uint16_t* NULL Number of audio channels (1 = mono, 2 = stereo).
bitrate uint32_t* NULL Audio bitrate in kbps from the source file tags/properties.

C representation: XBERGBBox is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGBBox does not appear anywhere in the generated header.

Bounding box in original image coordinates (x1, y1) top-left, (x2, y2) bottom-right.

Field Type Default Description
x1 float — Left edge (x-coordinate of the top-left corner).
y1 float — Top edge (y-coordinate of the top-left corner).
x2 float — Right edge (x-coordinate of the bottom-right corner).
y2 float — Bottom edge (y-coordinate of the bottom-right corner).

Since: v1.1

C representation: XBERGBedrockConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGBedrockConfig does not appear anywhere in the generated header.

AWS Bedrock configuration for bedrock/-prefixed models.

Mirrors liter-llm’s BedrockConfig. Every field is optional: anything left unset falls back to the standard AWS environment variables (AWS_DEFAULT_REGION / AWS_REGION, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN, BEDROCK_CROSS_REGION) and the default AWS credential chain. Leave the credential fields unset unless you have an explicit reason to pin them.

Debug is implemented by hand so the three credential fields are never printed.

Field Type Default Description
region const char* NULL AWS region (e.g. "us-east-1").
cross_region_prefix const char* NULL Cross-region inference profile prefix (e.g. "us").
access_key_id const char* NULL Explicit AWS access key ID. Secret — never logged.
secret_access_key const char* NULL Explicit AWS secret access key. Secret — never logged.
session_token const char* NULL Explicit AWS session token for temporary credentials. Secret — never logged.

C representation: XBERGBibtexMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGBibtexMetadata does not appear anywhere in the generated header.

BibTeX bibliography metadata.

Field Type Default Description
entry_count uintptr_t — Number of entries in the bibliography.
citation_keys const char* NULL BibTeX citation keys (e.g. "knuth1984") for all entries.
authors const char* NULL Author names collected across all bibliography entries.
year_range XBERGAlefHandle NULL Earliest and latest publication years found in the bibliography.
entry_types const char* NULL Count of entries grouped by BibTeX entry type (e.g. "article" → 5).

C representation: XBERGBoundingBox is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGBoundingBox does not appear anywhere in the generated header.

Bounding box coordinates for element positioning.

Field Type Default Description
x0 double — Left x-coordinate
y0 double — Bottom y-coordinate
x1 double — Right x-coordinate
y1 double — Top y-coordinate

C representation: XBERGBrowserConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGBrowserConfig does not appear anywhere in the generated header.

Browser fallback configuration.

Field Type Default Description
mode XBERGAlefHandle XBERG_AUTO When to use the headless browser fallback.
backend XBERGAlefHandle XBERG_CHROMIUMOXIDE Browser backend used to render JavaScript-heavy pages.
endpoint const char* NULL CDP WebSocket endpoint for connecting to an external browser instance.
timeout uint64_t 30000ms Timeout for browser page load and rendering (in milliseconds when serialized).
overall_timeout uint64_t 60000ms Overall deadline for a single browser fetch, covering browser launch (or page acquisition from a shared pool), page setup, navigation, rendering, and screenshot capture. Must exceed timeout to leave room for launch and setup overhead; a fetch that has not returned within this deadline fails with a timeout error (in milliseconds when serialized). Shutdown/teardown is governed separately by shutdown_timeout and is not counted against this deadline: an already-computed result is delivered to the caller without waiting for the browser process to exit.
shutdown_timeout uint64_t 5000ms How long to wait for the browser process to close and exit cleanly during teardown before the process is forcibly killed (in milliseconds when serialized).
wait XBERGAlefHandle XBERG_NETWORK_IDLE Wait strategy after browser navigation.
wait_selector const char* NULL CSS selector to wait for when wait is Selector.
extra_wait uint64_t* NULL Extra time to wait after the wait condition is met.
proxy XBERGAlefHandle NULL Proxy for browser fetches. Overrides CrawlConfig.proxy when set. Native backend supports http/https only (no SOCKS5).
block_url_patterns const char* NULL URL patterns to block before the network request fires. Supports * wildcards. Useful for skipping ads/analytics/large images. Honored by BrowserBackend.Native; chromiumoxide ignores this field today.
eval_script const char* NULL JavaScript snippet evaluated after navigation completes. Scraping captures the native backend result in ScrapeResult.browser.eval_result. Interactions run this script before page actions on both browser backends but do not include the script result in InteractionResult.
robots_user_agent const char* NULL User-agent used when fetching robots.txt. Defaults to BrowserConfig.user_agent (or crawlberg’s default) if unset. Native only.
capture_network_events int32_t false Capture the full network event stream into the result. Default false (only the document event is captured). Native only.
session_affinity int32_t true Enable session affinity: reuse chromiumoxide Pages for same-domain requests so cookies + fingerprint + solved challenges persist. Default: true. When false, each request gets a fresh Page.

C representation: XBERGCacheStats is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCacheStats does not appear anywhere in the generated header.

Aggregate statistics for a xberg cache directory.

Field Type Default Description
total_files uintptr_t — Total number of files currently in the cache directory.
total_size_mb double — Combined size of all cache files in megabytes.
available_space_mb double — Free disk space available on the cache volume, in megabytes.
oldest_file_age_days double — Age of the oldest cache file in days (0.0 if the cache is empty).
newest_file_age_days double — Age of the most recently written cache file in days (0.0 if the cache is empty).

Since: v1.0

C representation: XBERGCaptioningConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCaptioningConfig does not appear anywhere in the generated header.

Configuration for the VLM captioning post-processor.

Field Type Default Description
llm XBERGAlefHandle — LLM configuration used for the VLM call.
prompt const char* NULL Optional custom caption prompt. NULL uses the default RegionKind.Caption prompt that ships with crate.llm.region_extractor.
min_image_area uint32_t 1000 Skip images whose width * height is below this threshold (in pixels). Default 1_000 filters out icons and decorations.

C representation: XBERGCellChange is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCellChange does not appear anywhere in the generated header.

A single changed cell within a table.

Defined here (rather than only in crate.diff) so RevisionDelta can reference it unconditionally, without requiring the diff Cargo feature. crate.diff re-exports this type verbatim.

Field Type Default Description
row uintptr_t — Zero-based row index.
col uintptr_t — Zero-based column index.
from const char* — Value before the change.
to const char* — Value after the change.

C representation: XBERGChunk is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGChunk does not appear anywhere in the generated header.

A text chunk with optional embedding and metadata.

Chunks are created when chunking is enabled in ExtractionConfig. Each chunk contains the text content, optional embedding vector (if embedding generation is configured), and metadata about its position in the document.

Field Type Default Description
content const char* — The text content of this chunk.
chunk_type XBERGAlefHandle /* serde(default) */ Semantic structural classification of this chunk. Assigned by the heuristic classifier based on content patterns and heading context. Defaults to ChunkType.Unknown when no rule matches.
embedding const char* NULL Optional embedding vector for this chunk. Only populated when EmbeddingConfig is provided in chunking configuration. The dimensionality depends on the chosen embedding model.
sparse_embedding XBERGAlefHandle /* serde(default) */ Optional sparse (SPLADE) learned embedding for this chunk. Only populated when sparse-embedding generation is configured for chunking. NULL otherwise, including on builds without the sparse-embeddings feature. Uses the crate-root SparseEmbedding alias rather than crate.sparse_embeddings.SparseEmbedding directly: the sparse_embeddings module itself only compiles under sparse-embeddings/sparse-embedding-presets, while the crate-root alias is always defined (a field-compatible stub on builds without either feature), so this field — and Chunk itself — compiles on every feature combination, including the crate’s default features.
late_interaction XBERGAlefHandle /* serde(default) */ Optional ColBERT-style multi-vector (late-interaction) embedding for this chunk. Only populated when late-interaction embedding generation is configured for chunking. NULL otherwise, including on builds without the late-interaction feature. Uses the crate-root MultiVectorEmbedding alias for the same reason sparse_embedding uses SparseEmbedding — see that field’s docs.
metadata XBERGAlefHandle — Metadata about this chunk’s position and properties.

Since: v1.0

C representation: XBERGChunkClassificationConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGChunkClassificationConfig does not appear anywhere in the generated header.

Configuration for the chunk-classification post-processor.

Chunk classification is always multi-label: a chunk may match zero, one, or many of the configured definitions. This is the chunk-level equivalent of PageClassificationConfig, but scoped to individual chunks (ExtractedDocument.chunks) rather than whole pages, and built for large taxonomies where each label needs its own description rather than a bare name.

Field Type Default Description
prompt_template const char* NULL Minijinja prompt template. Receives {{ definitions }} (rendered label + description list) and {{ chunks }} (a numbered list of chunk texts in the current batch) variables. NULL lets the backend pick a sensible default.
definitions const char* — The set of label definitions the classifier may emit. Must contain at least one entry.
llm XBERGAlefHandle — LLM configuration used for classification.
batch_size uintptr_t 10 Number of chunks batched into a single LLM request. Larger batches amortize the fixed prompt cost (definitions block) across more chunks, at the risk of exceeding the model’s context window for very large taxonomies or chunk texts. Defaults to DEFAULT_BATCH_SIZE.
max_concurrency uintptr_t 4 Maximum number of in-flight batch requests. Bounds concurrency against the configured LLM provider. Defaults to DEFAULT_MAX_CONCURRENCY.

Since: v1.0

C representation: XBERGChunkClassificationDefinition is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGChunkClassificationDefinition does not appear anywhere in the generated header.

A single labeled definition the chunk classifier may emit.

Unlike PageClassificationConfig.labels (bare label names), chunk classification targets potentially large domain taxonomies where every label carries its own semantic description, letting the LLM disambiguate similarly named labels without relying on the label string alone.

Field Type Default Description
label const char* — Label name returned in ChunkMetadata.classifications.
description const char* — Semantic description of when this label applies. Injected verbatim into the classification prompt next to the label name.

C representation: XBERGChunkClassificationEnrichmentConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGChunkClassificationEnrichmentConfig does not appear anywhere in the generated header.

Chunk-classification enrichment knob: how to multi-label individual chunks.

Operates on ExtractedDocument.chunks in place — the caller must have already produced chunks (e.g. via ExtractionConfig.chunking) for this stage to have any effect; a document with no chunks is a no-op.

Field Type Default Description
config XBERGAlefHandle — Label-definition set and LLM/batching settings for the chunk-classification stage.

C representation: XBERGChunkInfo is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGChunkInfo does not appear anywhere in the generated header.

Information about a single chunk.

Field Type Default Description
index uint32_t — Zero-based chunk index.
pages XBERGAlefHandle — Page range for this chunk.
estimated_time_ms uint64_t — Estimated processing time for this chunk in milliseconds.

C representation: XBERGChunkMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGChunkMetadata does not appear anywhere in the generated header.

Metadata about a chunk’s position in the original document.

Field Type Default Description
byte_start uintptr_t — Byte offset where this chunk starts in the original text (UTF-8 valid boundary).
byte_end uintptr_t — Byte offset where this chunk ends in the original text (UTF-8 valid boundary).
token_count uintptr_t* NULL Number of tokens in this chunk (if available). This is calculated by the embedding model’s tokenizer if embeddings are enabled.
chunk_index uintptr_t — Zero-based index of this chunk in the document.
total_chunks uintptr_t — Total number of chunks in the document.
first_page uint32_t* NULL First page number this chunk spans (1-indexed). Only populated when page tracking is enabled in extraction configuration.
last_page uint32_t* NULL Last page number this chunk spans (1-indexed, equal to first_page for single-page chunks). Only populated when page tracking is enabled in extraction configuration.
heading_context XBERGAlefHandle /* serde(default) */ Heading context when using Markdown chunker. Contains the heading hierarchy this chunk falls under. Only populated when ChunkerType.Markdown is used.
heading_path const char* /* serde(default) */ Flattened heading trail from document root to this chunk’s section. Each element is a heading’s text, outermost first. Derived from heading_context when present; empty otherwise. Provides a binding-friendly, RAG-shaped breadcrumb without requiring callers to walk the nested HeadingContext structure.
image_indices const char* /* serde(default) */ Indices into ExtractedDocument.images for images on pages covered by this chunk. Contains zero-based indices into the top-level images collection for every image whose page_number falls within [first_page, last_page]. Empty when image extraction is disabled or the chunk spans no pages with images.
node_ids const char* /* serde(default) */ Ids of the DocumentNodes this chunk was derived from. Joins a chunk back to the structured document tree via DocumentNode.id. Populated from exact node provenance when available, with a textual containment fallback for rendered chunks that do not retain byte offsets.
page_spans const char* /* serde(default) */ Per-page bounding-box spans this chunk covers, for viewer highlighting (#1295). One entry per page the chunk overlaps, in page order — the first and last entries’ page fields equal first_page/last_page. Populated whenever page-boundary provenance is available (the same condition under which first_page/last_page are populated); each entry’s bbox is additionally populated when the document’s structured node tree (ExtractedDocument.document) is available, as the union of that page’s body-layer node bounding boxes found within this chunk. Empty when page-boundary provenance is unavailable (mirrors first_page/ last_page being NULL).
classifications const char* /* serde(default) */ Multi-label classification result for this chunk. Populated by the chunk-classification post-processor when ExtractionConfig.chunk_classification is set. A chunk may match zero, one, or many of the configured label definitions. Empty when chunk classification was not configured.

C representation: XBERGChunkingConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGChunkingConfig does not appear anywhere in the generated header.

Chunking configuration.

Configures text chunking for document content, including chunk size, overlap, trimming behavior, and optional embeddings.

Use ..the default constructor when constructing to allow for future field additions:

Field Type Default Description
max_characters uintptr_t 1000 Maximum size per chunk (in units determined by sizing). When sizing is Characters (default), this is the max character count. When using token-based sizing, this is the max token count. Default: 1000
overlap uintptr_t 200 Overlap between chunks (in units determined by sizing). Default: 200
trim int32_t true Whether to trim whitespace from chunk boundaries. Default: true
chunker_type XBERGAlefHandle XBERG_TEXT Type of chunker to use (Text or Markdown). Default: Text
embedding XBERGAlefHandle NULL Optional embedding configuration for chunk embeddings.
sparse_embedding XBERGAlefHandle NULL Optional sparse (SPLADE) embedding configuration for chunk embeddings. When set, sparse vectors are generated for each chunk’s content and attached via sparse_embedding. Requires the sparse-embeddings feature; without it, a warning is emitted and no sparse vectors are attached. Config-file only: like RerankerConfig and the local-ONNX branch of embedding, this has no CLI flag and no environment variable. Only the secret/identity fields of LLM-routed configs (model, API key, base URL) get that reach.
late_interaction XBERGAlefHandle NULL Optional late-interaction (ColBERT) embedding configuration for chunk embeddings. When set, multi-vector embeddings are generated for each chunk’s content and attached via late_interaction. Requires the late-interaction feature; without it, a warning is emitted and no late-interaction vectors are attached. Config-file only, for the same reason as sparse_embedding above.
preset const char* NULL Use a preset configuration (overrides individual settings if provided).
sizing XBERGAlefHandle XBERG_CHARACTERS How to measure chunk size. Default: Characters (Unicode character count). Enable chunking-tiktoken or chunking-tokenizers features for token-based sizing.
topic_threshold float* NULL Optional cosine similarity threshold for semantic topic boundary detection. Only used when chunker_type is Semantic and an EmbeddingConfig is provided. You almost never need to set this. When omitted, defaults to 0.75 which works well for most documents. Lower values detect more topic boundaries (more, smaller chunks); higher values detect fewer. Range: 0.0..=1.0.
table_chunking XBERGAlefHandle XBERG_SPLIT How to handle markdown tables that exceed the chunk size limit. Only applies when chunker_type is Markdown. - Split (default) — tables are split at row boundaries; continuation chunks do not repeat the header. - RepeatHeader — the table header row and separator are prepended to every continuation chunk so each chunk is self-contained. Default: Split

C representation: XBERGCitation is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCitation does not appear anywhere in the generated header.

A structured citation from a citation block.

Parsed from entries like: [^srcN]: source, locator, excerpt: "text"

Field Type Default Description
label const char* — The label of the citation (e.g., “src1” in [^src1]: ...).
source const char* — The source reference (path, URL, or identifier).
locator const char* NULL Optional locator within the source (e.g., “page 3” or “section 2.1”).
excerpt const char* NULL Optional excerpt — quoted text from the source.

C representation: XBERGCitationMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCitationMetadata does not appear anywhere in the generated header.

Citation file metadata (RIS, PubMed, EndNote).

Field Type Default Description
citation_count uintptr_t — Total number of citation records in the file.
format const char* NULL Detected citation file format (e.g. "ris", "pubmed", "endnote").
authors const char* NULL Author names collected across all citation records.
year_range XBERGAlefHandle NULL Earliest and latest publication years found in the file.
dois const char* NULL DOI identifiers found in the citation records.
keywords const char* NULL Keywords collected from all citation records.

C representation: XBERGClassificationLabel is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGClassificationLabel does not appear anywhere in the generated header.

A single label + confidence pair.

Field Type Default Description
label const char* — Label name as configured in PageClassificationConfig.labels.
confidence float* NULL Backend-reported confidence in [0.0, 1.0]. NULL when the backend (e.g. an LLM prompt without explicit confidence schema) did not report one.

C representation: XBERGCodeChunkInfo is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCodeChunkInfo does not appear anywhere in the generated header.

A single structurally-meaningful code chunk produced by tree-sitter parsing.

Purpose-built payload owned by xberg — deliberately does not expose the upstream tree_sitter_language_pack types, so binding generators never need to resolve an external crate’s types across FFI/language boundaries.

Field Type Default Description
text const char* — The raw source text of this chunk.
context_path const char* — Hierarchical path of enclosing structural items (e.g. ["MyClass", "my_method"]).
node_types const char* — Tree-sitter node kinds that appear at the top level of this chunk (e.g. "function_definition", "class_definition").
byte_start uintptr_t — Inclusive start byte offset of this chunk in the original source.
byte_end uintptr_t — Exclusive end byte offset of this chunk in the original source.

C representation: XBERGCodeDataAttribute is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCodeDataAttribute does not appear anywhere in the generated header.

An XML-style attribute attached to an Element node.

Populated only for CodeDataNodeKind.Element; always empty for KeyValue and Sequence nodes.

Field Type Default Description
name const char* — Attribute name (e.g. "class", "href").
value const char* — Attribute value as a raw string (quotes stripped).
byte_start uintptr_t — Inclusive start byte offset of the name="value" attribute token.
byte_end uintptr_t — Exclusive end byte offset of the name="value" attribute token.

C representation: XBERGCodeDataNode is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCodeDataNode does not appear anywhere in the generated header.

A node in the hierarchical data tree produced by data-format extraction.

Purpose-built payload owned by xberg — mirrors tree_sitter_language_pack.DataNode but flattens its Span down to plain byte offsets, so binding generators never need to resolve an external crate’s types across FFI/language boundaries.

Field Type Default Description
kind XBERGAlefHandle — Whether this node is a key/value pair, XML element, or sequence item.
key const char* /* serde(default) */ Key, attribute name, tag name, or positional index ("0", "1", …). NULL at the document root.
value const char* /* serde(default) */ Leaf scalar value, if any. NULL for containers (objects, arrays, XML elements with child elements).
attributes const char* /* serde(default) */ Attributes on element-shape nodes (XML STag attributes). Empty for all other kinds.
children const char* /* serde(default) */ Children for nested containers and XML element bodies.
byte_start uintptr_t — Inclusive start byte offset of this node in the original source.
byte_end uintptr_t — Exclusive end byte offset of this node in the original source.

C representation: XBERGCodeMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCodeMetadata does not appear anywhere in the generated header.

Code-format metadata: the structural chunks produced by tree-sitter parsing.

Wrapped by FormatMetadata.Code. Kept as a named struct (rather than an inline enum-variant body) so serde can tag it under internal tagging and utoipa can emit a referenceable CodeMetadata component in the OpenAPI schema.

Field Type Default Description
chunks const char* NULL Structural code chunks (function/class/module boundaries).
data XBERGAlefHandle NULL Hierarchical key/value data tree extracted from data-format source (JSON, YAML, TOML, XML, CSV, etc.), when data extraction was enabled.

C representation: XBERGConcurrencyConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGConcurrencyConfig does not appear anywhere in the generated header.

Controls thread usage for constrained environments.

Set max_threads to cap all internal thread pools (Rayon, ONNX Runtime intra-op), batch concurrency and Tesseract recognition to a single limit. Set max_concurrent_ocr to give recognition a limit of its own, which is the knob to reach for when the host has cores to spare but not the memory to run a recognition session on each of them. It is applied as given and is not capped by max_threads. The first extraction in a process fixes it for that process — see the field’s own documentation.

Without an explicit max_threads, the effective budget is min(detected_cpu_cores, 8) — a deliberate ceiling chosen for serverless/shared-tenant defaults, not a full-host auto-scale. On a host with more than 8 cores this means the extra cores go unused unless one of the following applies:

  • max_threads is set explicitly above 8 (the only way to exceed the ceiling on a bare-metal or VM host with no CPU quota).

  • The process runs under a Linux cgroup CPU quota (containers, Kubernetes resources.limits.cpu); in that case the quota itself is used as the ceiling instead of the hardcoded 8, since the quota already reflects a deliberately-configured resource limit.

When neither applies and the host has more than 8 cores, a single WARN-level log is emitted the first time the budget is resolved, naming the detected core count and the applied cap, so the ceiling is discoverable without reading source.

Field Type Default Description
max_threads uintptr_t* NULL Maximum number of threads for all internal thread pools. Caps Rayon global pool size, ONNX Runtime intra-op threads, and the combined document/inner-task budget for batch extraction. When NULL, the effective budget is min(detected_cpu_cores, 8) unless a Linux cgroup CPU quota is present, in which case the quota is used as the ceiling instead. On hosts with more than 8 cores and no cgroup quota, set max_threads explicitly to use the additional cores — the default will not scale past 8 on its own.
max_concurrent_ocr uintptr_t* NULL Maximum number of Tesseract recognition sessions that run at once. When NULL, recognition follows max_threads, reduced to the number of sessions the host’s free memory holds. Each session keeps its own page image and recognition working set resident, so a host with many cores and little memory needs this lower than the thread budget. Set it to 4 to keep the fixed limit that releases up to 1.2.6 applied. A value set here is applied as given: neither the thread budget nor the memory reading reduces it. Both of those bound the automatic limit, and a caller who names a number has already decided what the host can carry. The first extraction in a process fixes the limit for the rest of that process, and a later extraction that names a different value keeps the first one. The two limiters that enforce it — the admission semaphore in the Tesseract backend and the handle pool behind it — are built once inside a backend the plugin registry holds for the life of the process, and the pool’s capacity is fixed when it is constructed. A later value could therefore be reported but never enforced. Set it on the first extraction, or run one process per value.

C representation: XBERGContentConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGContentConfig does not appear anywhere in the generated header.

Content extraction and conversion configuration.

Controls how HTML is converted to the output format. Uses html-to-markdown-rs as the conversion engine for all formats (markdown, plain text, djot).

Field Type Default Description
output_format const char* "markdown" Output format: "markdown" (default), "plain", "djot".
preprocessing_preset const char* "standard" Preprocessing aggressiveness: "minimal", "standard" (default), "aggressive". - Minimal: only scripts/styles removed. - Standard: also removes nav, nav-hinted headers/footers/asides, forms. - Aggressive: removes all footers/asides unconditionally.
remove_navigation int32_t true Remove navigation elements (nav, breadcrumbs, menus). Default: true.
remove_forms int32_t true Remove form elements. Default: true.
strip_tags const char* NULL HTML tag names to strip (render children only, remove the tag wrapper). Default: [].
preserve_tags const char* NULL HTML tag names to preserve as raw HTML in output.
exclude_selectors const char* ["noscript"] CSS selectors for elements to exclude entirely (element + all content). Unlike strip_tags (which removes the wrapper but keeps children), excluded elements and all descendants are dropped. Supports CSS selectors: .class, #id, [attribute], compound selectors. Default: ["noscript"]. <noscript> fallback content (no-JS notices, tracking pixels, GTM iframes) is meant for browsers with JavaScript disabled, not for a markdown reader, and strip_tags cannot drop it — on preprocessing_preset: "standard" (crawlberg’s only path) it only removes the wrapper and still renders the children. Example: [".cookie-banner", "#ad-container", "[role='complementary']"]
skip_images int32_t false Skip image elements in output. Default: false.
max_depth uintptr_t* NULL Max DOM traversal depth. Prevents stack overflow on deeply nested HTML.
wrap int32_t false Enable line wrapping. Default: false.
wrap_width uintptr_t 80 Wrap width when wrap is enabled. Default: 80.
include_document_structure int32_t true Include document structure tree in output. Default: true.
extract_metadata int32_t true Prepend a YAML frontmatter block (title, description, etc., extracted from <head>) to the markdown output. Default: true. This only controls the frontmatter text inside markdown.content. Crawlberg never reads <head> metadata back out of the converter’s result – PageMetadata is populated independently by crate.html.metadata.extract_metadata from the parsed DOM, so turning this off does not lose any metadata field.

C representation: XBERGContentFilterConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGContentFilterConfig does not appear anywhere in the generated header.

Cross-extractor content filtering configuration.

Controls whether “furniture” content (headers, footers, page numbers, watermarks, repeating text) is included in or stripped from extraction results. Applies across all extractors (PDF, DOCX, RTF, ODT, HTML, etc.) with format-specific implementation.

When NULL on ExtractionConfig, each extractor uses its current default behavior unchanged.

Field Type Default Description
include_headers int32_t false Include running headers in extraction output. - PDF: Disables top-margin furniture stripping and prevents the layout model from treating PageHeader-classified regions as furniture. - DOCX: Includes document headers in text output. - RTF/ODT: Headers already included; this is a no-op when true. - HTML/EPUB: Keeps <header> element content. Default: false (headers are stripped or excluded).
include_footers int32_t false Include running footers in extraction output. - PDF: Disables bottom-margin furniture stripping and prevents the layout model from treating PageFooter-classified regions as furniture. - DOCX: Includes document footers in text output. - RTF/ODT: Footers already included; this is a no-op when true. - HTML/EPUB: Keeps <footer> element content. Default: false (footers are stripped or excluded).
include_footnotes int32_t false Include footnote bodies in extraction output. - PDF: Prevents the layout model from treating Footnote-classified regions as furniture, so footnote bodies survive alongside the main text instead of being silently dropped. - Other formats: No effect currently. Default: false (footnotes are stripped), matching the existing include_headers / include_footers defaults.
strip_repeating_text int32_t true Enable the heuristic cross-page repeating text detector. When true (default), text that repeats verbatim across a supermajority of pages is classified as furniture and stripped. Disable this if brand names or repeated headings are being incorrectly removed by the heuristic. This flag also gates a same-page rule: a body paragraph is removed when its text is also carried by a table detected on the same page (catches table content that PDF extraction renders both as a table and as body text). The comparison preserves case and never touches headings, list items, code blocks, formulas, or captions, so it cannot delete a body sentence merely for repeating an earlier heading’s words (GH#1623). Note: when a layout-detection model is active, the model may independently classify page-header / page-footer / footnote regions as furniture on a per-page basis. To preserve those regions, set include_headers = true, include_footers = true, include_footnotes = true, or any combination, in addition to disabling this flag. Primarily affects PDF extraction. Default: true.
include_watermarks int32_t false Include watermark text in extraction output. - PDF: Keeps watermark artifacts and arXiv identifiers. - Other formats: No effect currently. Default: false (watermarks are stripped).

C representation: XBERGContributorRole is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGContributorRole does not appear anywhere in the generated header.

JATS contributor with role.

Field Type Default Description
name const char* — Contributor display name.
role const char* NULL Contributor role (e.g. "author", "editor").

C representation: XBERGConversionOptions is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGConversionOptions does not appear anywhere in the generated header.

Main conversion options for HTML to Markdown conversion.

Use ConversionOptions.builder() to construct, or the default constructor for defaults.

Field Type Default Description
heading_style XBERGAlefHandle XBERG_ATX Heading style to use in Markdown output (ATX # or Setext underline).
list_indent_type XBERGAlefHandle XBERG_SPACES How to indent nested list items (spaces or tab).
list_indent_width uintptr_t 2 Number of spaces (or tabs) to use for each level of list indentation.
bullets const char* "-*+" Bullet character(s) to use for unordered list items (e.g. "-", "*").
strong_em_symbol const char* "*" Character used for bold/italic emphasis markers (* or _).
escape_asterisks int32_t false Escape * characters in plain text to avoid unintended bold/italic.
escape_underscores int32_t false Escape _ characters in plain text to avoid unintended bold/italic.
escape_misc int32_t false Escape miscellaneous Markdown metacharacters ([]()# etc.) in plain text.
escape_ascii int32_t false Escape ASCII characters that have special meaning in certain Markdown dialects.
code_language const char* "" Default language annotation for fenced code blocks that have no language hint.
autolinks int32_t true Automatically convert bare URLs into Markdown autolinks.
default_title int32_t false Emit a default title when no <title> tag is present.
br_in_tables int32_t false Render <br> elements inside table cells as literal line breaks.
compact_tables int32_t false Emit tables without column padding (compact GFM format). When true, column widths are not computed and cells are emitted with no trailing spaces. Separator rows use exactly --- per column. Produces token-efficient output suitable for RAG / LLM contexts. Default false (aligned padding preserved).
highlight_style XBERGAlefHandle XBERG_DOUBLE_EQUAL Style used for <mark> / highlighted text (e.g. ==text==).
extract_metadata int32_t true Populate result.metadata with <head> / <meta> extraction (title, description, Open Graph, Twitter Card, JSON-LD, …). Default true. Disabling skips the metadata pass only — table extraction into result.tables runs unconditionally.
whitespace_mode XBERGAlefHandle XBERG_NORMALIZED Controls how whitespace sequences are normalised in the converted output. - WhitespaceMode.Normalized (default) — collapses consecutive whitespace characters (spaces, tabs, newlines) to a single space, matching browser rendering behaviour. - WhitespaceMode.Strict — preserves all whitespace exactly as it appears in the source HTML, including runs of spaces and embedded newlines. Choose Strict only when the source HTML uses deliberate whitespace (e.g. pre-formatted content outside <pre> tags). For most documents Normalized produces cleaner output.
strip_newlines int32_t false Strip all newlines from the output, producing a single-line result.
wrap int32_t false Wrap long lines at wrap_width characters.
wrap_width uintptr_t 80 Maximum output line width in characters when wrap is true (default 80). Lines are broken at word boundaries so that no line exceeds this length. A value of 0 is treated as “no limit” — equivalent to leaving wrap disabled. Has no effect when wrap is false.
convert_as_inline int32_t false Treat the entire document as inline content (no block-level wrappers).
sub_symbol const char* "" Markdown notation for subscript text (e.g. "~").
sup_symbol const char* "" Markdown notation for superscript text (e.g. "^").
newline_style XBERGAlefHandle XBERG_SPACES How to encode hard line breaks (<br>) in Markdown.
code_block_style XBERGAlefHandle XBERG_BACKTICKS Style used for fenced code blocks (backticks or tilde).
keep_inline_images_in const char* NULL HTML tag names whose <img> children are kept inline instead of block.
preprocessing XBERGAlefHandle — Options for the HTML pre-processing pass applied before conversion begins. Pre-processing runs before the HTML is handed to the converter and can perform operations such as unwrapping redundant wrapper elements, removing tracking pixels, and normalising vendor-specific markup. See PreprocessingOptions for the full set of knobs. Defaults to the standard preprocessing options, which enables the standard cleaning passes. Set individual fields on PreprocessingOptions (or construct via ConversionOptions.builder) to opt in or out of specific passes.
encoding const char* "utf-8" Expected character encoding of the input HTML (default "utf-8").
debug int32_t false Emit debug information during conversion.
strip_tags const char* NULL HTML tag names whose content is stripped from the output entirely.
preserve_tags const char* NULL HTML tag names that are preserved verbatim in the output.
skip_images int32_t false Skip conversion of <img> elements (omit images from output).
url_escape_style XBERGAlefHandle XBERG_ANGLE URL encoding strategy for link and image destinations. Controls how special characters in URL destinations are escaped: - UrlEscapeStyle.Angle (default) — wraps the destination in angle brackets when it contains spaces or newlines. Some parsers misinterpret > inside such a destination. - UrlEscapeStyle.Percent — percent-encodes every character that is not an RFC 3986 unreserved character or /, producing a destination that all Markdown parsers handle correctly even when the URL contains <, >, spaces, or parentheses.
link_style XBERGAlefHandle XBERG_INLINE Link rendering style (inline or reference).
output_format XBERGAlefHandle XBERG_MARKDOWN Target output format (Markdown, plain text, etc.).
include_document_structure int32_t false Include structured document tree in result.
extract_images int32_t false Extract inline images from data URIs and SVGs.
max_image_size uint64_t 5242880 Maximum decoded image size in bytes (default 5MB).
capture_svg int32_t false Capture SVG elements as images.
infer_dimensions int32_t true Infer image dimensions from data.
max_depth uintptr_t* NULL Maximum DOM traversal depth. NULL uses the library’s internal native-stack safety limit. Explicit values above that safety limit are clamped to prevent process-aborting stack overflows on pathologically deep DOM trees.
exclude_selectors const char* NULL CSS selectors for elements to exclude entirely (element + all content). Unlike strip_tags (which removes the tag wrapper but keeps children), excluded elements and all their descendants are dropped from the output. Supports any CSS selector that tl supports: tag names, .class, #id, [attribute], etc. Invalid selectors are silently skipped at conversion time. Example: [".cookie-banner", "#ad-container", "[role='complementary']"]
tier_strategy XBERGAlefHandle XBERG_AUTO Which conversion tier to use. - TierStrategy.Auto (default) — automatically choose the best path. - TierStrategy.Tier2 — always use the Tier-2 DOM-walk path. - TierStrategy.Tier1 — always attempt Tier-1 (testkit only).
base_url const char* NULL Base URL to resolve relative href/src destinations against. When set, every relative link and image/media destination (a href, img src and its lazy-load fallbacks, srcset, graphic url/href/xlink:href/src, iframe/audio/video/source src) is resolved to an absolute URL before being written to the Markdown output, so the result is followable without the reader knowing where the source HTML came from. A <base href> in the document, if present, is honored the way a browser honors it: it is itself resolved against base_url, and that combined result becomes the effective base every other relative reference resolves against. Non-hierarchical schemes (mailto:, tel:, javascript:, data:) and already-absolute URLs are left unchanged. An unset (default) or unparseable base_url, and a relative reference that fails to resolve, leave the original attribute text unchanged – this option never panics and never corrupts a destination it cannot confidently resolve. Default NULL — no resolution, output byte-identical to versions before this option existed.

C representation: XBERGCoreProperties is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCoreProperties does not appear anywhere in the generated header.

Dublin Core metadata from docProps/core.xml

Contains standard metadata fields defined by the Dublin Core standard and Office-specific extensions.

Field Type Default Description
title const char* NULL Document title
subject const char* NULL Document subject/topic
creator const char* NULL Document creator/author
keywords const char* NULL Keywords or tags
description const char* NULL Document description/abstract
last_modified_by const char* NULL User who last modified the document
revision const char* NULL Revision number
created const char* NULL Creation timestamp (ISO 8601)
modified const char* NULL Last modification timestamp (ISO 8601)
category const char* NULL Document category
content_status const char* NULL Content status (Draft, Final, etc.)
language const char* NULL Document language
identifier const char* NULL Unique identifier
version const char* NULL Document version
last_printed const char* NULL Last print timestamp (ISO 8601)

C representation: XBERGCrawlConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCrawlConfig does not appear anywhere in the generated header.

Configuration for crawl, scrape, and map operations.

Field Type Default Description
max_depth uintptr_t* NULL Maximum crawl depth (number of link hops from the start URL).
max_pages uintptr_t* NULL Maximum number of pages to crawl.
max_links_per_page uintptr_t* NULL Maximum links enqueued from a single page. Defaults to 10000. Bounds the work one hostile or pathological page can create; links past the cap are dropped and a warning is logged.
max_concurrent uintptr_t* NULL Maximum number of concurrent requests.
crawl_strategy XBERGAlefHandle XBERG_BFS Traversal order. Defaults to breadth-first. A frontier or strategy set explicitly on CrawlEngineBuilder takes precedence over this field.
content_filter XBERGAlefHandle NULL Content filter applied to each page. NULL keeps every page. A content filter set explicitly on CrawlEngineBuilder takes precedence.
bm25_query const char* NULL Query the BM25 content filter scores pages against. Required by ContentFilterKind.Bm25.
bm25_threshold double* NULL Minimum BM25 score a page must reach to be kept. Defaults to 0.0.
respect_robots_txt int32_t false Whether to respect robots.txt directives. A crawl that respects them also honours the page’s own robots instructions: it does not follow the links of a page marked nofollow by its robots meta tag or an X-Robots-Tag header. A link marked rel="nofollow" is a hint, not a robots directive, and is still followed. A noindex page is still crawled and its links followed; the page result marks it with noindex_detected.
soft_http_errors int32_t false When true, HTTP-level error responses (404 NotFound, 403 Forbidden, WAF blocks) are surfaced as ScrapeResult records with the matching status_code rather than raised as CrawlError. Default false preserves the historical throw-on-error contract for direct fetches. Independently of this flag, 404s reached at the end of a redirect chain are always surfaced softly — the user opted into redirect-following, so receiving a 404 there is part of the normal flow rather than an unexpected error.
user_agent const char* NULL Custom user-agent string.
stay_on_domain int32_t false Whether to confine document links (.pdf, .docx, .zip, …) to the seed domain. Page links are always confined to the seed host, widened to its subdomains by Self.allow_subdomains; this flag does not loosen that. It applies only to document links, which are classified by file extension before their host is considered and so are followed cross-host by default – the usual case being documents served from a CDN or object store. Set this to true to require documents to live on the seed domain too.
allow_subdomains int32_t false Whether subdomains of the seed host are in scope. Applies to page links unconditionally, and to document links when Self.stay_on_domain is set.
include_paths const char* NULL Regex patterns for paths to include during crawling.
exclude_paths const char* NULL Regex patterns for paths to exclude during crawling.
path_patterns_match_query int32_t false Whether include_paths/exclude_paths match against path?query instead of just path. Defaults to false, matching path only: a pattern anchored with $ (e.g. /feed/?$) changes meaning once the query joins the matched text, so this must stay opt-in rather than silently changing what an existing config matches.
dedup_include_query int32_t false Whether the crawl-dedup key includes the (sorted) query string. Defaults to false, matching historical behavior: /item?id=1 and /item?id=2 are treated as one page and only the first is fetched. true keeps the query, sorted, in the key, so each distinct query is fetched once.
strip_tracking_params int32_t false Whether to strip tracking_params from a discovered URL before it is deduplicated, fetched, and reported. Defaults to false, so no tracking parameters are stripped unless explicitly enabled.
tracking_params const char* ["utm_*", "fbclid", "gclid", "ref"] Query parameter name patterns to strip when strip_tracking_params is true. A pattern ending in * matches by prefix (utm_* matches utm_source, utm_campaign, …); any other pattern matches the parameter name exactly. Defaults to ["utm_*", "fbclid", "gclid", "ref"], applied only once strip_tracking_params is enabled.
custom_headers const char* NULL Custom HTTP headers to send with each request.
request_timeout uint64_t 30000ms Timeout for individual HTTP requests (in milliseconds when serialized).
rate_limit_ms uint64_t* NULL Per-domain rate limit in milliseconds. When set, enforces a minimum delay between requests to the same domain. Defaults to 200ms when NULL.
max_redirects uintptr_t 10 Maximum number of redirects to follow.
retry_count uintptr_t 0 Number of retry attempts for failed requests. Bounded by MAX_RETRY_COUNT.
retry_codes const char* NULL HTTP status codes that should trigger a retry. When empty, every rate limit, server error, bad gateway and timeout is retried. When set, only a failure whose status is listed is retried, so a timeout without a response is not.
retry_initial_delay_ms uint64_t 100 Initial delay, in milliseconds, before the first retry. Doubled on each subsequent attempt (capped at retry_max_delay_ms). Defaults to 100ms.
retry_max_delay_ms uint64_t 60000 Upper bound, in milliseconds, on the exponential retry backoff. Defaults to 60s.
rate_limit_jitter_ratio double 0 Fraction of the per-domain rate-limit delay to randomly jitter by, in [0.0, 1.0]. 0.0 (the default) applies no jitter and preserves the previous fixed-interval behaviour; 0.1 jitters the delay by up to ±10%.
cookies_enabled int32_t false Whether to enable cookie handling.
auth XBERGAlefHandle NULL Authentication configuration.
max_body_size uintptr_t* NULL Maximum response body size in bytes. NULL does not mean unbounded: an unset cap falls back to a 100 MiB safety ceiling, because HTTP responses are decompressed while being read and a few hundred compressed bytes can otherwise expand to gigabytes in memory. To read bodies larger than that, set this explicitly.
remove_tags const char* NULL CSS selectors for tags to remove from HTML before processing.
content XBERGAlefHandle — Content extraction and conversion configuration.
map_limit uintptr_t* NULL Maximum number of URLs to return from a map operation.
map_search const char* NULL Search filter for map results (case-insensitive substring match on URLs).
download_assets int32_t false Whether to download assets (CSS, JS, images, etc.) from the page.
asset_types const char* NULL Filter for asset categories to download.
max_asset_size uintptr_t* NULL Maximum size in bytes for individual asset downloads.
browser XBERGAlefHandle — Browser configuration.
proxy XBERGAlefHandle NULL Proxy configuration for HTTP requests.
user_agents const char* NULL List of user-agent strings for rotation. If non-empty, overrides user_agent.
capture_screenshot int32_t false Whether to capture a screenshot when using the browser. Only supported by scrape() with BrowserBackend.Chromiumoxide and BrowserMode.Always or Stealth. A screenshot is 100–500 KB of PNG per page, so crawl() does not carry screenshots in CrawlPageResult/CrawlResult at all — a multi-thousand-page crawl holding one per page in memory is not a safe default. Setting this with any other configuration (a different backend, BrowserMode.Auto/Never, or during crawl()) has no effect and logs a warning rather than silently doing nothing.
follow_document_urls int32_t false Re-enqueue discovered LinkType.Document URLs into the crawl frontier so the crawl follows links from document pages (PDFs, etc.) as it would from HTML pages. Default: false (documents terminate at materialisation).
document_url_depth uint32_t* NULL Maximum document-depth (from the seed URL through document links only) when follow_document_urls is true. NULL means inherit max_depth. Independent of max_depth: a document URL is enqueued only if BOTH the outer max_depth and (if set) document_url_depth permit it.
download_documents int32_t true Whether to download non-HTML documents (PDF, DOCX, images, code, etc.) instead of skipping them. Defaults to true — unlike download_assets and capture_screenshot, which default to false.
document_max_size uintptr_t* 52428800 Maximum size in bytes for document downloads. Defaults to 50 MB.
document_mime_types const char* NULL Allowlist of MIME types to download. If empty, uses built-in defaults.
document_output_dir const char* NULL Directory to stream downloaded document bytes into instead of holding them in memory on DownloadedDocument.content. When set, content is left empty and DownloadedDocument.content_path is populated with <dir>/<content_hash>.<ext>. NULL (default) preserves today’s in-memory-only behavior. Has no effect on wasm32, which has no filesystem — use document_content_encoding there instead.
document_content_encoding XBERGAlefHandle NULL Opt-in encoding that duplicates DownloadedDocument.content into a serializable field for language bindings that need the bytes in-memory (content itself is alef(skip)ed). NULL (default) means no encoding is produced. Independent of document_output_dir — set both to get a file on disk and an in-memory copy.
warc_output const char* NULL Path to write WARC output. If NULL, WARC output is disabled.
browser_profile const char* NULL Named browser profile for persistent sessions (cookies, localStorage). Chromiumoxide backend only. The native backend runs an in-process JavaScript engine with no Chrome process and therefore no profile directory, so this is ignored there and logs a warning. It is also ignored — with a warning — when a shared browser pool is in use (the pool launches before any per-crawl config exists) or when connecting to an external CDP endpoint whose process crawlberg does not own.
save_browser_profile int32_t false Whether to save changes back to the browser profile on exit.
ssrf XBERGAlefHandle crawlberg::SsrfPolicy::from_env() SSRF policy for outbound network requests. Default: deny private networks, allow http/https only, max 5 redirects. All policy fields are exposed to language bindings. wasm32 (including Node.js): deny_private does not stop hostname-based requests. There is no DNS resolution on this target, so only a literal IP host is checked against the policy — a domain name is always permitted, regardless of deny_private. Under Node, where fetch enforces no CORS, this means a service embedding the wasm binding can be driven to internal hosts by domain name even with deny_private = true. Enforce egress restrictions at the network layer for that deployment target; do not rely on this field. See crawlberg.net.validate_url.
ssrf_deny_private_explicit int32_t* NULL Pins SsrfPolicy.deny_private to a caller-chosen value, bypassing the CRAWLBERG_ALLOW_PRIVATE_NETWORK operator override entirely for this config. ssrf.deny_private is a plain, always-serialized bool: several alef-generated bindings construct SsrfPolicy.default() (hardcoding deny_private: true) whenever their caller never touches SSRF settings at all, so true on that field alone cannot distinguish “the caller wants private networks denied” from “the binding’s own structural default landed on true”. The environment variable exists precisely to resolve that ambiguity in the common case by treating any true as inconclusive and deferring to the operator. Set this field when that default-deferral is wrong for your call — e.g. a test that must prove deny_private: true still denies even while the operator has set CRAWLBERG_ALLOW_PRIVATE_NETWORK suite-wide for every other call. NULL (default) preserves today’s behavior: the environment variable may still flip ssrf.deny_private to false. Some(value) pins ssrf.deny_private to value and the environment variable is not consulted for this config.

Since: v1.1

C representation: XBERGCsvConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCsvConfig does not appear anywhere in the generated header.

Configuration for CSV/TSV extraction.

When unset (ExtractionConfig.csv == None), the extractor keeps its existing default behavior: the delimiter is auto-detected by sampling the file (comma, tab, pipe, or semicolon), and no line is treated as a comment.

Field Type Default Description
delimiter const char* NULL Field delimiter, as a single-character string (e.g. ",", ";", "\t", "|"). When NULL (default), the delimiter is auto-detected from a sample of the file. Must be exactly one ASCII byte when set — ExtractionConfig.validate rejects an empty string or a multi-byte value with a helpful error. The TSV MIME type (text/tab-separated-values) always forces \t regardless of this setting.
comment_prefixes const char* NULL Line prefixes that mark a comment line to skip entirely during row parsing (e.g. ["#"]). A line is treated as a comment when its trimmed start matches any of these prefixes exactly. Default: empty, meaning no line is treated as a comment (matches the pre-existing extractor behavior).

C representation: XBERGCsvMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGCsvMetadata does not appear anywhere in the generated header.

CSV/TSV file metadata.

Field Type Default Description
row_count uint32_t — Total number of data rows (excluding the header row if present).
column_count uint32_t — Number of columns detected.
delimiter const char* NULL Field delimiter character (e.g. "," or "\t").
has_header int32_t — Whether the first row was treated as a header.
column_types const char* NULL Inferred data type for each column (e.g. "string", "integer", "float").

C representation: XBERGDbfFieldInfo is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDbfFieldInfo does not appear anywhere in the generated header.

dBASE field information.

Field Type Default Description
name const char* — Field (column) name.
field_type const char* — dBASE field type character (e.g. "C" for character, "N" for numeric).

C representation: XBERGDbfMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDbfMetadata does not appear anywhere in the generated header.

dBASE (DBF) file metadata.

Field Type Default Description
record_count uintptr_t — Total number of data records in the DBF file.
field_count uintptr_t — Number of field (column) definitions.
fields const char* NULL Descriptor for each field in the table schema.

C representation: XBERGDetectResponse is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDetectResponse does not appear anywhere in the generated header.

MIME type detection response.

Field Type Default Description
mime_type const char* — Detected MIME type
filename const char* NULL Original filename (if provided)

C representation: XBERGDetectionResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDetectionResult does not appear anywhere in the generated header.

Page-level detection result containing all detections and page metadata.

Field Type Default Description
page_width uint32_t — Page width in pixels (as seen by the model).
page_height uint32_t — Page height in pixels (as seen by the model).
detections const char* — All layout detections on this page after postprocessing.

C representation: XBERGDiffHunk is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDiffHunk does not appear anywhere in the generated header.

A single contiguous hunk in a unified diff.

Field Type Default Description
from_line uintptr_t — Starting line number in the old content (0-indexed).
from_count uintptr_t — Number of lines from the old content in this hunk.
to_line uintptr_t — Starting line number in the new content (0-indexed).
to_count uintptr_t — Number of lines from the new content in this hunk.
lines const char* — Lines that make up this hunk.

C representation: XBERGDiffOptions is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDiffOptions does not appear anywhere in the generated header.

Options controlling how two ExtractedDocument values are compared.

Field Type Default Description
include_metadata int32_t true Include metadata changes in the diff. Default: true.
include_embedded int32_t true Include embedded-children changes in the diff. Default: true.
max_content_chars uintptr_t* NULL Truncate content to this many characters before diffing. Useful for very large documents where only the first N characters matter. NULL means no truncation.

C representation: XBERGDjotAttributeGroup is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDjotAttributeGroup does not appear anywhere in the generated header.

Attributes associated with a named Djot element.

Field Type Default Description
identifier const char* — Element identifier used by the Djot attribute map.
attributes XBERGAlefHandle — Attributes associated with the element.

C representation: XBERGDjotContent is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDjotContent does not appear anywhere in the generated header.

Comprehensive Djot document structure with semantic preservation.

This type captures the full richness of Djot markup, including:

  • Block-level structures (headings, lists, blockquotes, code blocks, etc.)
  • Inline formatting (emphasis, strong, highlight, subscript, superscript, etc.)
  • Attributes (classes, IDs, key-value pairs)
  • Links, images, footnotes
  • Math expressions (inline and display)
  • Tables with full structure

Available when the djot feature is enabled.

Field Type Default Description
plain_text const char* — Plain text representation for backwards compatibility
blocks const char* — Structured block-level content
metadata XBERGAlefHandle — Metadata from YAML frontmatter
tables const char* — Extracted tables as structured data
images const char* — Extracted images with metadata
links const char* — Extracted links with URLs
footnotes const char* — Footnote definitions
attributes const char* /* serde(default) */ Attributes mapped by element identifier (if present)

C representation: XBERGDjotImage is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDjotImage does not appear anywhere in the generated header.

Image element in Djot.

Field Type Default Description
src const char* — Image source URL or path
alt const char* — Alternative text
title const char* NULL Optional title
attributes XBERGAlefHandle NULL Element attributes

C representation: XBERGDjotLink is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDjotLink does not appear anywhere in the generated header.

Link element in Djot.

Field Type Default Description
url const char* — Link URL
text const char* — Link text content
title const char* NULL Optional title
attributes XBERGAlefHandle NULL Element attributes

C representation: XBERGDoctorCheck is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDoctorCheck does not appear anywhere in the generated header.

A single doctor verdict: what was checked, the outcome, and why.

Field Type Default Description
name const char* — Check identifier, e.g. ocr.tesseract or layout.rtdetr.
status XBERGAlefHandle — Pass / warn / fail / skip verdict.
message const char* — One-line reason or detail (e.g. missing language, resolved path, error).

C representation: XBERGDoctorReport is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDoctorReport does not appear anywhere in the generated header.

Aggregate doctor report over all configured backends and settings.

Field Type Default Description
checks const char* NULL Individual check verdicts, in execution order.

C representation: XBERGDocumentBoundary is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentBoundary does not appear anywhere in the generated header.

Detected document boundary within a PDF.

Field Type Default Description
start_page uint32_t — 1-indexed start page (inclusive).
end_page uint32_t — 1-indexed end page (inclusive).
confidence float — Confidence in this boundary, [0.0, 1.0].
reason XBERGAlefHandle — Reason for the boundary detection.

C representation: XBERGDocumentCounts is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentCounts does not appear anywhere in the generated header.

Cheap structural counts for an extracted document.

Populated on every ExtractedDocument returned by extract / extract_batch, regardless of whether the heavy pages / images collections are materialized. A caller that only needs “how many pages / tables / images did this document have?” (reporting, cost estimation, progress, quotas) can read these without enabling per-page or per-image extraction.

The page count comes from the parse (the extractor already walks the page tree); it does not require opting into per-page content. pages is 0 for inputs that are not page-addressable (e.g. plain text).

Field Type Default Description
pages uintptr_t — Total pages in the source document (0 when not page-addressable).
tables uintptr_t — Tables detected in the document.
images uintptr_t — Images detected in the document.

C representation: XBERGDocumentExtractor is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentExtractor does not appear anywhere in the generated header.

Trait for document extractor plugins.

Implement this trait to add support for new document formats or override built-in extraction behavior. Foreign-language bindings expose the DocumentExtractor.extract method, which accepts ExtractInput and returns an ExtractedDocument.

When multiple extractors support the same MIME type, the registry selects the extractor with the highest priority value. Use this to:

  • Override built-in extractors (priority > 50)
  • Provide fallback extractors (priority < 50)
  • Implement specialized extractors for specific use cases

Default priority is 50.

Extractors must be thread-safe (Send + Sync) to support concurrent extraction.

Binding-safe extraction entry point for foreign-language plugin bridges.

Accepts the same unified input shape as the public extraction API and returns one extracted document result.

Signature:

XBERGAlefHandle xberg_document_extractor_extract(XBERGAlefHandle this, XBERGAlefHandle input, XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_document_extractor_extract(instance, 0, 0);

Parameters:

Name Type Required Description
input XBERGAlefHandle Yes The input data
config XBERGAlefHandle Yes The configuration options

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.

xberg_document_extractor_supported_mime_types()
Section titled “xberg_document_extractor_supported_mime_types()”

Get the list of MIME types supported by this extractor.

Can include exact MIME types and prefix patterns:

  • Exact: "application/pdf", "text/plain"
  • Prefix: "image/*" (matches any image type)

Returns:

A slice of MIME type strings.

Signature:

const char* xberg_document_extractor_supported_mime_types(XBERGAlefHandle this);

Example:

const char* result = xberg_document_extractor_supported_mime_types(instance);

Returns: const char*

Get the priority of this extractor.

Higher priority extractors are preferred when multiple extractors support the same MIME type.

  • 0-25: Fallback/low-quality extractors
  • 26-49: Alternative extractors
  • 50: Default priority (built-in extractors)
  • 51-75: Premium/enhanced extractors
  • 76-100: Specialized/high-priority extractors

Returns:

Priority value (default: 50)

Signature:

int32_t xberg_document_extractor_priority(XBERGAlefHandle this);

Example:

int32_t result = xberg_document_extractor_priority(instance);

Returns: int32_t

Optional: Check if this extractor can handle a specific file.

Allows for more sophisticated detection beyond MIME types. Defaults to true (rely on MIME type matching).

Returns:

true if the extractor can handle this file, false otherwise.

Signature:

int32_t xberg_document_extractor_can_handle(XBERGAlefHandle this, const char* path, const char* mime_type);

Example:

int32_t result = xberg_document_extractor_can_handle(instance, "value", "value");

Parameters:

Name Type Required Description
path const char* Yes The path
mime_type const char* Yes The mime type

Returns: int32_t


C representation: XBERGDocumentMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentMetadata does not appear anywhere in the generated header.

Metadata about a document for analysis.

Field Type Default Description
mime_type const char* — MIME type of the document.
size_bytes uint64_t — File size in bytes.
page_count uint32_t* NULL Page count (if known, e.g., from previous analysis).
force_ocr int32_t — Whether OCR is forced regardless of text layer.
user_chunk_config XBERGAlefHandle NULL User-provided chunk configuration overrides.
chunking_enabled int32_t — Whether chunking is enabled for this job.

C representation: XBERGDocumentNode is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentNode does not appear anywhere in the generated header.

A single node in the document tree.

Each node has deterministic id, typed content, optional parent/children for tree structure, and metadata like page number, bounding box, and content layer.

Field Type Default Description
id const char* /* serde(default) */ Deterministic identifier (hash of node type + text + page + position). Stable and unique within a single extraction response: every internal construction path threads the node’s position (its index in DocumentStructure.nodes) into the hash, so identical (node_type, text, page) tuples at different positions never collide. Always serialised — ChunkMetadata.node_ids references it to join chunks back to the nodes they were derived from. #[serde(default)] covers the missing-field case on inbound JSON (e.g. documents serialised before this field existed).
content XBERGAlefHandle — Node content — tagged enum, type-specific data only.
parent uint32_t* NULL Parent node index (NULL = root-level node).
children const char* /* serde(default) */ Child node indices in reading order.
content_layer XBERGAlefHandle /* serde(default) */ Content layer classification. Always serialised — Kotlin-Android (and any other typed binding) treats the field as non-nullable, so omitting it from the JSON wire would break consumer deserialisation. #[serde(default)] covers the missing-field case on inbound JSON.
page uint32_t* NULL Page number where this node starts (1-indexed).
page_end uint32_t* NULL Page number where this node ends (for multi-page tables/sections).
bbox XBERGAlefHandle NULL Bounding box in document coordinates.
annotations const char* /* serde(default) */ Inline annotations (formatting, links) on this node’s text content. Only meaningful for text-carrying nodes; empty for containers.
attributes const char* NULL Format-specific key-value attributes. Extensible bag for miscellaneous data without a dedicated typed field: CSS classes, LaTeX environment names, Excel cell formulas, slide layout names, etc.

C representation: XBERGDocumentRelationship is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentRelationship does not appear anywhere in the generated header.

A resolved relationship between two nodes in the document tree.

Field Type Default Description
source uint32_t — Source node index (the referencing node).
target uint32_t — Target node index (the referenced node).
kind XBERGAlefHandle — Semantic kind of the relationship.

C representation: XBERGDocumentRevision is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentRevision does not appear anywhere in the generated header.

A single tracked change embedded in a document.

Populated by per-format extractors that understand change-tracking metadata (DOCX w:ins/w:del/w:rPrChange, ODT text:change-*, …). Every extractor defaults to ExtractedDocument.revisions = None until a format-specific implementation is added.

Field Type Default Description
revision_id const char* — Format-specific revision identifier. For DOCX this is the w:id attribute value on the change element (e.g. "42"). When the attribute is absent a synthetic fallback is generated ("docx-ins-0", "docx-del-3", …).
author const char* NULL Display name of the author who made this change, when available.
timestamp const char* NULL ISO-8601 timestamp of the change, when available. Stored as a plain string so this type remains FFI-friendly and unconditionally available without the chrono optional dep. DOCX populates this from the w:date attribute (e.g. "2024-03-15T10:30:00Z").
kind XBERGAlefHandle — Semantic kind of this revision.
anchor XBERGAlefHandle NULL Best-effort document location for this revision. Resolution is format-dependent and may be NULL when the location cannot be determined (e.g. changes inside table cells before table-cell anchor support is added).
delta XBERGAlefHandle — The content changes that make up this revision.

C representation: XBERGDocumentStructure is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentStructure does not appear anywhere in the generated header.

Top-level structured document representation.

A flat array of nodes with index-based parent/child references forming a tree. Root-level nodes have parent: None. Use body_roots() and furniture_roots() to iterate over top-level content by layer.

Call validate() after construction to verify all node indices are in bounds and parent-child relationships are bidirectionally consistent.

Field Type Default Description
nodes const char* NULL All nodes in document/reading order.
source_format const char* NULL Origin format identifier (e.g. “docx”, “pptx”, “html”, “pdf”). Allows renderers to apply format-aware heuristics when converting the document tree to output formats.
relationships const char* NULL Resolved relationships between nodes (footnote refs, citations, anchor links, etc.). Populated during derivation from the internal document representation. Empty when no relationships are detected.
node_types const char* NULL Sorted, deduplicated list of node type names present in this document. Each value is the snake_case node_type tag of the corresponding NodeContent variant (e.g. "paragraph", "heading", "table", …). Computed from nodes via DocumentStructure.finalize_node_types. Empty until that method is called (internal construction paths call it at the end of derivation).

C representation: XBERGDocumentSummary is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocumentSummary does not appear anywhere in the generated header.

Summary of an extracted document.

Field Type Default Description
text const char* — Summary text (plain prose).
strategy XBERGAlefHandle — Strategy that produced this summary.
token_count uint32_t* NULL Approximate token count of the summary, when known.

C representation: XBERGDocxAppProperties is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocxAppProperties does not appear anywhere in the generated header.

Application properties from docProps/app.xml for DOCX

Contains Word-specific document statistics and metadata.

Field Type Default Description
application const char* NULL Application name (e.g., “Microsoft Office Word”)
app_version const char* NULL Application version
template const char* NULL Template filename
total_time int32_t* NULL Total editing time in minutes
pages int32_t* NULL Number of pages
words int32_t* NULL Number of words
characters int32_t* NULL Number of characters (excluding spaces)
characters_with_spaces int32_t* NULL Number of characters (including spaces)
lines int32_t* NULL Number of lines
paragraphs int32_t* NULL Number of paragraphs
company const char* NULL Company name
doc_security int32_t* NULL Document security level
scale_crop int32_t* NULL Scale crop flag
links_up_to_date int32_t* NULL Links up to date flag
shared_doc int32_t* NULL Shared document flag
hyperlinks_changed int32_t* NULL Hyperlinks changed flag

C representation: XBERGDocxMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGDocxMetadata does not appear anywhere in the generated header.

Word document metadata.

Extracted from DOCX files using shared Office Open XML metadata extraction. Integrates with office_metadata module for core/app/custom properties.

Field Type Default Description
core_properties XBERGAlefHandle NULL Core properties from docProps/core.xml (Dublin Core metadata) Contains title, creator, subject, keywords, dates, etc. Shared format across DOCX/PPTX/XLSX documents.
app_properties XBERGAlefHandle NULL Application properties from docProps/app.xml (Word-specific statistics) Contains word count, page count, paragraph count, editing time, etc. DOCX-specific variant of Office application properties.
custom_properties const char* NULL Custom properties from docProps/custom.xml (user-defined properties) Contains key-value pairs defined by users or applications. Values can be strings, numbers, booleans, or dates.

C representation: XBERGElement is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGElement does not appear anywhere in the generated header.

Semantic element extracted from document.

Represents a logical unit of content with semantic classification, unique identifier, and metadata for tracking origin and position.

Field Type Default Description
element_id const char* /* serde(default) */ Deterministic element identifier. Empty only when deserializing legacy payloads that predate this field’s wire representation.
element_type XBERGAlefHandle — Semantic type of this element
text const char* — Text content of the element
metadata XBERGAlefHandle — Metadata about the element

C representation: XBERGElementMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGElementMetadata does not appear anywhere in the generated header.

Metadata for a semantic element.

Field Type Default Description
page_number uint32_t* NULL Page number (1-indexed)
filename const char* NULL Source filename or document name
coordinates XBERGAlefHandle NULL Bounding box coordinates if available
element_index uintptr_t* NULL Position index in the element sequence
additional const char* — Additional custom metadata

C representation: XBERGEmailAttachment is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmailAttachment does not appear anywhere in the generated header.

Email attachment representation.

Contains metadata and optionally the content of an email attachment.

Field Type Default Description
name const char* NULL Attachment name (from Content-Disposition header)
filename const char* NULL Filename of the attachment
mime_type const char* NULL MIME type of the attachment
size uintptr_t* NULL Size in bytes
is_image int32_t — Whether this attachment is an image
data const uint8_t* NULL Attachment data (if extracted). Uses bytes.Bytes for cheap cloning of large buffers.

C representation: XBERGEmailConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmailConfig does not appear anywhere in the generated header.

Configuration for email extraction.

Field Type Default Description
msg_fallback_codepage uint32_t* NULL Windows codepage number to use when an MSG file contains no codepage property. Defaults to NULL, which falls back to windows-1252. If an unrecognized or invalid codepage number is supplied (including 0), the behavior silently falls back to windows-1252 — the same as when the MSG file itself contains an unrecognized codepage. No error or warning is emitted. Users should verify output when supplying unusual values. Common values: - 1250: Central European (Polish, Czech, Hungarian, etc.) - 1251: Cyrillic (Russian, Ukrainian, Bulgarian, etc.) - 1252: Western European (default) - 1253: Greek - 1254: Turkish - 1255: Hebrew - 1256: Arabic - 932: Japanese (Shift-JIS) - 936: Simplified Chinese (GBK)

C representation: XBERGEmailExtractionResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmailExtractionResult does not appear anywhere in the generated header.

Email extraction result.

Complete representation of an extracted email message (.eml or .msg) including headers, body content, and attachments.

Field Type Default Description
subject const char* NULL Email subject line
from_email const char* NULL Sender email address
to_emails const char* — Primary recipient email addresses
cc_emails const char* — CC recipient email addresses
bcc_emails const char* — BCC recipient email addresses
date const char* NULL Email date/timestamp
message_id const char* NULL Message-ID header value
plain_text const char* NULL Plain text version of the email body
html_content const char* NULL HTML version of the email body
content const char* — Cleaned/processed text content. Aliased as cleaned_text for back-compat.
attachments const char* — List of email attachments
metadata const char* — Additional email headers and metadata

C representation: XBERGEmailMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmailMetadata does not appear anywhere in the generated header.

Email metadata extracted from .eml and .msg files.

Includes sender/recipient information, message ID, and attachment list.

Field Type Default Description
from_email const char* NULL Sender’s email address
from_name const char* NULL Sender’s display name
to_emails const char* NULL Primary recipients
cc_emails const char* NULL CC recipients
bcc_emails const char* NULL BCC recipients
message_id const char* NULL Message-ID header value
attachments const char* NULL List of attachment filenames

C representation: XBERGEmbeddedChanges is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmbeddedChanges does not appear anywhere in the generated header.

Changes to embedded archive children between two results.

Field Type Default Description
added const char* NULL Children present in b but not in a (matched by path).
removed const char* NULL Children present in a but not in b (matched by path).
changed const char* NULL Children present in both but with differing content (matched by path). Each entry holds the diff of the nested ExtractedDocument.

C representation: XBERGEmbeddedDiff is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmbeddedDiff does not appear anywhere in the generated header.

Diff for a single embedded archive entry that appears in both results.

Field Type Default Description
path const char* — Archive-relative path identifying this entry.
diff XBERGAlefHandle — The recursive diff of the entry’s extraction result.

C representation: XBERGEmbeddedFile is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmbeddedFile does not appear anywhere in the generated header.

Embedded file descriptor extracted from the PDF name tree.

Field Type Default Description
name const char* — The filename as stored in the PDF name tree.
data const uint8_t* — Raw file bytes from the embedded stream (already decompressed by lopdf).
compressed_size uintptr_t — Compressed byte count of the original stream (before decompression). Used by callers to compute the decompression ratio and detect zip-bomb-style attacks that embed a tiny compressed stream expanding to gigabytes of data.
mime_type const char* NULL MIME type if specified in the filespec, otherwise NULL.

C representation: XBERGEmbeddingBackend is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmbeddingBackend does not appear anywhere in the generated header.

Trait for in-process embedding backend plugins.

Async to match the convention used by other plugin hooks such as OcrBackend and PostProcessor. Host-language bridges (PyO3, napi-rs, Rustler, extendr, magnus, ext-php-rs, C FFI, etc.) wrap their synchronous host callables in spawn_blocking or the equivalent to satisfy the async signature.

Backends must be Send + Sync + 'static. They are stored in Arc<dyn EmbeddingBackend> and called concurrently from xberg’s chunking pipeline. If the backend’s underlying model isn’t thread-safe, the backend itself must serialize access internally (e.g. via Mutex<Inner>).

  • embed(texts) MUST return exactly texts.len() vectors, each of length self.dimensions(). The dispatcher in crate.embeddings.embed_texts validates this before returning to downstream consumers; a non-conforming backend surfaces as a XbergError.Validation, not a panic.

  • embed may be called from any thread. Its future must be Send (enforced by async_trait when #[async_trait] is used on non-WASM targets).

  • dimensions() is called exactly once at registration, immediately after initialize() succeeds. The returned value is cached by the registry and used for all subsequent shape validation. Lazy-loading implementations can defer model loading into initialize() and report the real dimension afterwards. Later mutations of the backend’s reported dimension are not observed by xberg — implementations that need to change dimension must unregister and re-register.

  • shutdown() (inherited from Plugin) may be invoked concurrently with an in-flight embed() call. Implementations must tolerate this — e.g. by letting in-flight calls finish using resources held via the Arc<dyn EmbeddingBackend> reference, and only releasing shared state that isn’t needed by embed.

The synchronous embed_texts entry uses tokio.task.block_in_place to await the trait’s async embed, which requires a multi-thread tokio runtime. Callers running inside a current_thread runtime (e.g. #[tokio.test] without flavor = "multi_thread", or tokio.runtime.Builder.new_current_thread()) must use embed_texts_async instead, which awaits directly without block_in_place.

Embedding vector dimension. Must be > 0 and must match the length of every vector returned by embed.

Signature:

uintptr_t xberg_embedding_backend_dimensions(XBERGAlefHandle this);

Example:

uintptr_t result = xberg_embedding_backend_dimensions(instance);

Returns: uintptr_t

Embed a batch of texts, returning one vector per input in order.

Errors:

Implementations should return Plugin for backend-specific failures. The dispatcher layers its own validation (length, per-vector dimension) on top.

Signature:

const char* xberg_embedding_backend_embed(XBERGAlefHandle this, const char* texts);

Example:

const char* result = xberg_embedding_backend_embed(instance, NULL);

Parameters:

Name Type Required Description
texts const char* Yes The texts

Returns: const char*

Errors: Returns NULL on error.


C representation: XBERGEmbeddingConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEmbeddingConfig does not appear anywhere in the generated header.

Embedding configuration for text chunks.

Configures embedding generation using ONNX models via the vendored embedding engine. Requires the embeddings feature to be enabled.

Field Type Default Description
model XBERGAlefHandle XBERG_PRESET { name: "balanced" } The embedding model to use (defaults to “gte-modernbert-base” preset if not specified)
normalize int32_t true Whether to normalize embedding vectors (recommended for cosine similarity)
batch_size uintptr_t 32 Batch size for embedding generation
show_download_progress int32_t false Show model download progress. When enabled, transfer progress for the model, tokenizer and config files is reported at info level on the xberg.model_download target while they download (#279). Covers both local backends (ONNX and static/model2vec). A warm Hugging Face cache transfers nothing and so reports nothing. Ignored by EmbeddingModelType.Llm and EmbeddingModelType.Plugin, which download no model.
cache_dir const char* NULL Optional alternate Hugging Face cache root for model files. When unset, hf-hub follows HF_HUB_CACHE, HUGGINGFACE_HUB_CACHE, HF_HOME, XDG, and platform defaults. Prefer those environment variables when configuring the cache process-wide.
acceleration XBERGAlefHandle NULL Hardware acceleration for the embedding ONNX model. When set, controls which execution provider (CPU, CUDA, CoreML, TensorRT) is used for inference. Defaults to NULL (auto-select per platform).
max_embed_duration_secs uint64_t* 60 Maximum wall-clock duration (in seconds) for a single embed() call when using EmbeddingModelType.Plugin. Applies only to the in-process plugin path — protects against hung host-language backends (e.g. a Python callback deadlocked on the GIL, a model stuck on CUDA OOM retries, etc.). On timeout, the dispatcher returns Plugin instead of blocking forever. NULL disables the timeout. The default (60 seconds) is conservative for common in-process inference; increase for large batches on slow hardware.
max_sequence_length uintptr_t* NULL Maximum number of tokens fed to the tokenizer before truncation when embedding a chunk with a local ONNX model (Preset/Custom). A chunk longer than this many tokens has its tail dropped before inference, so only the prefix contributes to the stored vector. NULL falls back to 512 (the historical default). The effective value is always capped at the model’s own model_max_length, so raising it past what the model supports has no effect — set it to match a long-context model (e.g. 8192 for Jina/Nomic) so long chunks embed in full. Ignored by the Llm and Plugin model types, which own their own tokenization.

C representation: XBERGEntity is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEntity does not appear anywhere in the generated header.

A single named entity detected in the extracted text.

Field Type Default Description
category XBERGAlefHandle — Canonical category the entity belongs to (PERSON, ORG, LOCATION, etc.).
text const char* — Raw mention text exactly as it appeared in the source.
start uint32_t — Byte-offset span in ExtractedDocument.content where the mention starts.
end uint32_t — Byte-offset span in ExtractedDocument.content where the mention ends (exclusive).
confidence float* NULL Backend-reported confidence in [0.0, 1.0]. NULL when the backend does not expose confidence scores.

C representation: XBERGEpubMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGEpubMetadata does not appear anywhere in the generated header.

EPUB metadata (Dublin Core extensions).

Field Type Default Description
coverage const char* NULL Dublin Core coverage field (geographic or temporal scope).
dc_format const char* NULL Dublin Core format field (media type of the resource).
relation const char* NULL Dublin Core relation field (related resource identifier).
source const char* NULL Dublin Core source field (origin resource identifier).
dc_type const char* NULL Dublin Core type field (nature or genre of the resource).
cover_image const char* NULL Path or identifier of the cover image within the EPUB container.

C representation: XBERGErrorMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGErrorMetadata does not appear anywhere in the generated header.

Error metadata (for batch operations).

Field Type Default Description
error_type const char* — Machine-readable error type identifier (e.g. “UnsupportedFormat”).
message const char* — Human-readable error description.

C representation: XBERGExcelMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExcelMetadata does not appear anywhere in the generated header.

Excel/spreadsheet format metadata.

Identifies the document as a spreadsheet source via the FormatMetadata.Excel discriminant. Sheet count and sheet names are stored inside this struct.

Field Type Default Description
sheet_count uint32_t* NULL Number of sheets in the workbook.
sheet_names const char* NULL Names of all sheets in the workbook.

C representation: XBERGExcelSheet is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExcelSheet does not appear anywhere in the generated header.

Single Excel worksheet.

Represents one sheet from an Excel workbook with its content converted to Markdown format and dimensional statistics.

Field Type Default Description
name const char* — Sheet name as it appears in Excel
markdown const char* — Sheet content converted to Markdown tables
row_count uintptr_t — Number of rows
col_count uintptr_t — Number of columns
cell_count uintptr_t — Total number of non-empty cells
table_cells const char* NULL Pre-extracted table cells (2D vector of cell values) Populated during markdown generation to avoid re-parsing markdown. None for empty sheets.

C representation: XBERGExcelWorkbook is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExcelWorkbook does not appear anywhere in the generated header.

Excel workbook representation.

Contains all sheets from an Excel file (.xlsx, .xls, etc.) with extracted content and metadata.

Field Type Default Description
sheets const char* — All sheets in the workbook
metadata const char* — Workbook-level metadata (author, creation date, etc.)
revisions const char* /* serde(default) */ Collaborative-edit revision headers from xl/revisions/revisionHeaders.xml. Populated for legacy shared-workbook .xlsx files that contain the xl/revisions/ directory. Each <header> element maps to one DocumentRevision { kind: FormatChange } carrying the header’s guid (→ revision_id), userName (→ author), and dateTime (→ timestamp). anchor and delta are NULL/empty for v1 (per-cell log parsing is a follow-up). NULL when xl/revisions/revisionHeaders.xml is absent.

C representation: XBERGExtractInput is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractInput does not appear anywhere in the generated header.

Unified extraction input for all public extraction entry points.

Field Type Default Description
kind XBERGAlefHandle XBERG_URI Source kind. bytes requires bytes; uri requires uri.
bytes const uint8_t* NULL Raw bytes for kind = "bytes".
uri const char* NULL Local path, file:// URI, or HTTP(S) URL for kind = "uri".
mime_type const char* NULL MIME type hint.
filename const char* NULL Filename hint used for MIME detection and metadata.
config XBERGAlefHandle NULL Per-input extraction overrides.

C representation: XBERGExtractedDocument is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractedDocument does not appear anywhere in the generated header.

Document extracted by the core extraction pipeline.

extract and extract_batch return an ExtractionResult envelope whose results field contains these per-document payloads.

Field Type Default Description
content const char* — Plain-text representation of the extracted document content.
mime_type const char* — MIME type of the source document (e.g. "application/pdf").
metadata XBERGAlefHandle — Document-level metadata (author, title, dates, format-specific fields).
extraction_method XBERGAlefHandle NULL Extraction strategy used to produce the returned text. Populated when the extractor can reliably distinguish native text extraction, OCR-only extraction, or mixed native/OCR output.
tables const char* NULL Tables extracted from the document, each with structured cell data.
counts XBERGAlefHandle — Cheap structural counts (pages, tables, images). Always populated by the extraction pipeline, even when the pages / images collections are NULL. See DocumentCounts.
detected_languages const char* NULL ISO 639-1 language codes detected in the document content.
detected_language_confidences const char* NULL Structured per-language detection results: confidence, document share, script, and reliability, alongside the ISO-code-only detected_languages (#261). One entry per language in detected_languages, in the same order. NULL under the same conditions as detected_languages: detection disabled, empty input text, or no language met the configured min_confidence.
chunks const char* NULL Text chunks when chunking is enabled. When chunking configuration is provided, the content is split into overlapping chunks for efficient processing. Each chunk contains the text, optional embeddings (if enabled), and metadata about its position.
images const char* NULL Extracted images from the document. When image extraction is enabled via ImageExtractionConfig, this field contains all images found in the document with their raw data and metadata. Each image may optionally contain a nested ocr_result if OCR was performed.
pages const char* NULL Per-page content when page extraction is enabled. When page extraction is configured, the document is split into per-page content with tables and images mapped to their respective pages.
elements const char* NULL Semantic elements when element-based result format is enabled. When result_format is set to ElementBased, this field contains semantic elements with type classification, unique identifiers, and metadata for Unstructured-compatible element-based processing.
djot_content XBERGAlefHandle NULL Rich Djot content structure (when extracting Djot documents). When extracting Djot documents with structured extraction enabled, this field contains the full semantic structure including: - Block-level elements with nesting - Inline formatting with attributes - Links, images, footnotes - Math expressions - Complete attribute information The content field still contains plain text for backward compatibility. Always NULL for non-Djot documents.
ocr_elements const char* NULL OCR elements with full spatial and confidence metadata. When OCR is performed with element extraction enabled, this field contains the structured representation of detected text including: - Bounding geometry (rectangles or quadrilaterals) - Confidence scores (detection and recognition) - Rotation information - Hierarchical relationships (Tesseract only) This field preserves all metadata that would otherwise be lost when converting to plain text or markdown output formats. Only populated when OcrElementConfig.include_elements is true.
document XBERGAlefHandle NULL Structured document tree (when document structure extraction is enabled). When include_document_structure is true in ExtractionConfig, this field contains the full hierarchical representation of the document including: - Heading-driven section nesting - Table grids with cell-level metadata - Content layer classification (body, header, footer, footnote) - Inline text annotations (formatting, links) - Bounding boxes and page numbers Independent of result_format — can be combined with Unified or ElementBased.
extracted_keywords const char* NULL Extracted keywords when keyword extraction is enabled. When keyword extraction (RAKE or YAKE) is configured, this field contains the extracted keywords with scores, algorithm info, and position data. Previously stored in metadata.additional["keywords"].
quality_score double* NULL Text cleanliness/readability score from quality analysis. A value between 0.0 and 1.0 describing the quality of the text that was retained. This is not a completeness or recall score: clean text can score highly even when an extractor omitted or rejected other content. Inspect processing_warnings separately for known degraded or partial extraction. When the text came from OCR and the result carries enough recognized words to judge, this score is additionally capped by the mean OCR recognition confidence. Text that looks clean but that OCR itself had little confidence in therefore cannot score high. A native, non-OCR extraction is not capped. Previously stored in metadata.additional["quality_score"].
processing_warnings const char* NULL Non-fatal warnings collected during processing pipeline stages. Captures errors from optional pipeline features (embedding, chunking, language detection, output formatting) that don’t prevent extraction but may indicate degraded or incomplete results. These warnings are independent of quality_score, which assesses only retained text. Previously stored as individual keys in metadata.additional.
annotations const char* NULL PDF annotations extracted from the document. When annotation extraction is enabled via PdfConfig.extract_annotations, this field contains text notes, highlights, links, stamps, and other annotations found in PDF documents.
children const char* NULL Nested extraction results from archive contents. When extracting archives, each processable file inside produces its own full extraction result. Set to NULL for non-archive formats. Use max_archive_depth in config to control recursion depth.
uris const char* NULL URIs/links discovered during document extraction. Contains hyperlinks, image references, citations, email addresses, and other URI-like references found in the document. Always extracted when present in the source document.
revisions const char* NULL Tracked changes embedded in the source document. Populated by per-format extractors that understand change-tracking metadata (DOCX w:ins/w:del/w:rPrChange, ODT text:change-*, …). Every extractor defaults to NULL until its format-specific implementation is added. Extractors that do populate this field follow the “accepted-changes” convention: inserted text is present in content, deleted text is absent — the revision list is the separate audit trail.
structured_output const char* NULL Structured extraction output from LLM-based JSON schema extraction. When structured_extraction is configured in ExtractionConfig, the extracted document content is sent to a VLM with the provided JSON schema. The response is parsed and stored here as a JSON value matching the schema.
code_intelligence const char* NULL Code intelligence results from tree-sitter analysis. Populated when extracting source code files with the tree-sitter feature. Contains metrics, structural analysis, imports/exports, comments, docstrings, symbols, diagnostics, and optionally chunked code segments. Stored as an opaque JSON value so that all language bindings (Go, Java, C#, …) can deserialize it as a raw JSON object rather than a typed struct. The underlying type is tree_sitter_language_pack.ProcessResult.
llm_usage const char* NULL LLM token usage and cost data for all LLM calls made during this extraction. Contains one entry per LLM call. Multiple entries are produced when VLM OCR, structured extraction, or LLM embeddings run during the same extraction. NULL when no LLM was used.
entities const char* NULL Named entities detected in content by the NER post-processor. NULL when no NER backend is configured. Populated by the xberg-gliner ONNX backend or the LLM-driven backend (see crates/xberg/src/text/ner/).
summary XBERGAlefHandle NULL Summary of content produced by the summarisation post-processor. NULL when summarisation is not configured. Populated by the TextRank extractive backend (deterministic, no external service) or by the liter-llm-driven abstractive backend.
extraction_confidence XBERGAlefHandle NULL Confidence score computed by the heuristics pipeline. Populated when the heuristics feature is enabled and confidence scoring has been performed. Combines text-coverage, OCR aggregate confidence, and schema-compliance into a single [0, 1] value. NULL when confidence scoring is not configured or the feature is absent.
translation XBERGAlefHandle NULL Translation of content produced by the translation post-processor. NULL when translation is not configured.
page_classifications const char* NULL Per-page classifications produced by the page-classification post-processor. NULL when classification is not configured.
redaction_report XBERGAlefHandle NULL Audit report of redactions applied by the redaction post-processor. The redaction processor rewrites content, formatted_content, every chunk’s text, and the textual fields of entities / summary / translation / page_classifications in place. This report describes what was found and how it was replaced. NULL when redaction is not configured.
formulas const char* NULL Mathematical formulas recognized in the document. Populated from every source that produces formulas: layout-guided OCR (with geometry), VLM OCR (text only), and markup extraction (DOCX, PPTX, ODT, EPUB, HTML, JATS, LaTeX, Markdown, and related formats, without geometry). Empty when the document contains no formulas.
form_fields const char* NULL Form fields extracted from a PDF’s AcroForm or XFA structure. Populated by the PDF extractor when PdfConfig.extract_form_fields is enabled (default) and the document is a fillable form. Empty otherwise.

C representation: XBERGExtractedImage is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractedImage does not appear anywhere in the generated header.

Extracted image from a document.

Contains raw image data, metadata, and optional nested OCR results. Raw bytes allow cross-language compatibility - users can convert to PIL.Image (Python), Sharp (Node.js), or other formats as needed.

Field Type Default Description
data const uint8_t* — Raw image data (PNG, JPEG, WebP, etc. bytes). Uses bytes.Bytes for cheap cloning of large buffers.
format const char* — Image format (e.g., “jpeg”, “png”, “webp”) Uses Cow<’static, str> to avoid allocation for static literals.
image_index uint32_t — Zero-indexed position of this image in the document/page
page_number uint32_t* NULL Page/slide number where image was found (1-indexed)
width uint32_t* NULL Image width in pixels
height uint32_t* NULL Image height in pixels
colorspace const char* NULL Colorspace information (e.g., “RGB”, “CMYK”, “Gray”)
bits_per_component uint32_t* NULL Bits per color component (e.g., 8, 16)
is_mask int32_t — Whether this image is a mask image
description const char* NULL Optional description of the image
ocr_result XBERGAlefHandle NULL Nested OCR extraction result (if image was OCRed) When OCR is performed on this image, the result is embedded here rather than in a separate collection, making the relationship explicit.
bounding_box XBERGAlefHandle NULL Bounding box of the image on the page (PDF coordinates: x0=left, y0=bottom, x1=right, y1=top). Only populated for PDF-extracted images when position data is available from the PDF extractor.
source_path const char* NULL Original source path of the image within the document archive (e.g., “media/image1.png” in DOCX). Used for rendering image references when the binary data is not extracted.
image_kind XBERGAlefHandle NULL Heuristic classification of what this image likely depicts. NULL if classification was disabled or inconclusive.
kind_confidence float* NULL Confidence score for image_kind, in the range 0.0 to 1.0.
cluster_id uint32_t* NULL Identifier shared across images that form a single logical figure (e.g. all raster tiles of one technical drawing). NULL for singletons.
caption const char* NULL VLM-generated caption describing the image, when captioning is configured. Populated by the captioning post-processor (crates/xberg/src/plugins/processor/builtin/captioning.rs), which routes each image through crate.llm.region_extractor.extract_region_with_vlm in caption mode. NULL when captioning is disabled or the VLM declined to caption.
qr_codes const char* NULL QR codes decoded from this image, when QR detection is enabled. Populated by the QR post-processor (crates/xberg/src/extractors/qr.rs) via the pure-Rust rqrr decoder. NULL when QR detection is disabled; an empty Some([]) when detection ran but found nothing.
data_base64 const char* NULL Base64-encoded copy of data; populated when ImageExtractionConfig.include_data_base64 is true. Omitted from JSON by default; use instead of data in JSON-only clients.

C representation: XBERGExtractedUri is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractedUri does not appear anywhere in the generated header.

A URI extracted from a document.

Represents any link, reference, or resource pointer found during extraction. The kind field classifies the URI semantically, while label carries optional human-readable display text.

Field Type Default Description
url const char* — The URL or path string.
label const char* NULL Optional display text / label for the link.
page uint32_t* NULL Optional page number where the URI was found (1-indexed).
kind XBERGAlefHandle — Semantic classification of the URI.

C representation: XBERGExtractionConfidence is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractionConfidence does not appear anywhere in the generated header.

Combined confidence on [0, 1].

When OCR did not run, the ocr_aggregate weight folds into text_coverage so the weighted sum still totals 1.0.

Field Type Default Description
text_coverage float — Fraction of pages with a usable text layer.
ocr_aggregate float* NULL OCR recognition confidence, word-count-weighted across every recognized word, when OCR ran; NULL when it did not.
schema_compliance XBERGAlefHandle — Whether the merged output validates against the preset schema.
combined float — Weighted blend in [0, 1]. The value compared against the fallback threshold.

C representation: XBERGExtractionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractionConfig does not appear anywhere in the generated header.

Main extraction configuration.

This struct contains all configuration options for the extraction process. It can be loaded from TOML, YAML, or JSON files, or created programmatically.

Field Type Default Description
mime_detection_policy XBERGAlefHandle XBERG_PREFER_CONTENT Controls whether MIME inference prefers content, a supported extension, or content alone.
use_cache int32_t true Enable caching of extraction results
enable_quality_processing int32_t true Enable quality post-processing
ocr XBERGAlefHandle NULL OCR configuration. NULL does not run OCR for documents that already have usable text. Under OcrStrategy.Auto, a PDF with no text layer at all (a scan) is still routed to OCR with default settings so it is not returned empty (#1338). Set Self.disable_ocr to hard-disable OCR regardless of the detected content.
force_ocr int32_t false Force OCR even for searchable PDFs
ocr_strategy XBERGAlefHandle XBERG_AUTO Which pages get OCR’d when neither force_ocr nor force_ocr_pages applies. Defaults to OcrStrategy.Auto, which OCRs only pages whose native text fails a quality check. Only applies to PDF documents. Cannot be OcrStrategy.ScannedPages while disable_ocr is true.
force_ocr_pages const char* NULL Force OCR on specific pages only (1-indexed page numbers, must be >= 1). When set, only the listed pages are OCR’d regardless of text layer quality. Unlisted pages use native text extraction. Ignored when force_ocr is true. Only applies to PDF documents. Duplicates are automatically deduplicated. An ocr config is recommended for backend/language selection; defaults are used if absent.
disable_ocr int32_t false Disable OCR entirely, even for images. When true, OCR is skipped for all document types. Images return metadata only (dimensions, format, EXIF) without text extraction. PDFs use only native text extraction without OCR fallback. Cannot be true simultaneously with force_ocr.
ocr_near_empty_fallback int32_t* NULL Whether a PDF whose native text layer is empty or near-empty falls back to OCR under OcrStrategy.Auto (GH#1752). NULL (the default) derives the answer from whether Self.ocr is set, which is exactly the behaviour shipped before this field existed: with an ocr block the fallback always runs, and without one it runs only when the page’s native text is completely empty (#1338). That carve-out means a scanned page carrying a visible page label, Bates number or scanner stamp keeps only that label. Some(true) applies the near-empty fallback without requiring an ocr block: the text-quality gate decides, using OcrConfig.default’s thresholds when no block is present, and an automatic OCR backend must be registered. Some(false) suppresses the fallback even when an ocr block is present.
ocr_scanned_page_quality_gate int32_t* NULL Whether OcrStrategy.ScannedPages folds the per-page text-quality gate into its page selection, on top of the pages scan detection flagged (GH#1752). NULL (the default) derives the answer from whether Self.ocr is set, which is exactly the behaviour shipped before this field existed. Without an ocr block ScannedPages therefore degrades to detected-scans-only. Some(true) runs the gate without requiring an ocr block, using OcrConfig.default’s thresholds; an automatic OCR backend must be registered. Some(false) selects detected scans only even when an ocr block is present, which is the setting that avoids paying for recognition across a whole mixed document.
ocr_embedded_images int32_t* NULL Whether images embedded in a container document (DOCX, PPTX, ODT, HTML, …) are sent to OCR (GH#1752). NULL (the default) derives the answer from whether Self.ocr is set, which is exactly the behaviour shipped before this field existed. Some(true) recognises picture text without requiring an ocr block; Some(false) suppresses it even when a block is present. Orthogonal to ImageExtractionConfig.run_ocr_on_images, which still has to be true (its own default) for embedded-image OCR to run. See Self.runs_ocr_on_embedded_images.
chunking XBERGAlefHandle NULL Text chunking configuration (None = chunking disabled)
content_filter XBERGAlefHandle NULL Content filtering configuration (None = use extractor defaults). Controls whether document “furniture” (headers, footers, watermarks, repeating text) is included in or stripped from extraction results. See ContentFilterConfig for per-field documentation.
images XBERGAlefHandle NULL Image extraction configuration (None = no image extraction)
pdf_options XBERGAlefHandle NULL PDF-specific options (None = use defaults)
token_reduction XBERGAlefHandle NULL Token reduction configuration (None = no token reduction)
language_detection XBERGAlefHandle NULL Language detection configuration (None = no language detection)
pages XBERGAlefHandle NULL Page extraction configuration (None = no page tracking)
keywords XBERGAlefHandle NULL Keyword extraction configuration (None = no keyword extraction)
postprocessor XBERGAlefHandle NULL Post-processor configuration (None = use defaults)
html_options XBERGAlefHandle NULL HTML to Markdown conversion options (None = use defaults) Configure how HTML documents are converted to Markdown, including heading styles, list formatting, code block styles, and preprocessing options.
html_output XBERGAlefHandle NULL Styled HTML output configuration. When set alongside output_format = OutputFormat.Html, the extraction pipeline uses StyledHtmlRenderer which emits stable kb-* CSS class hooks on every structural element and optionally embeds theme CSS or user-supplied CSS in a <style> block. When NULL, the existing plain comrak-based HTML renderer is used.
extraction_timeout_secs uint64_t* 600 Default per-file timeout in seconds for batch extraction. When set, each file in a batch will be canceled after this duration unless overridden by FileExtractionConfig.timeout_secs. Defaults to Some(600) (10 minutes) to prevent pathological files (e.g. deeply nested archives, documents with millions of cells) from running indefinitely and exhausting caller resources, while still giving slow paths (VLM-based OCR, large scanned documents) enough headroom to finish. Set to NULL to disable the timeout for trusted input or long-running workloads.
max_concurrent_extractions uintptr_t* NULL Maximum concurrent document extractions in batch operations. This is a ceiling within the configured total thread budget, not an independent pool size. When unset, the scheduler derives document and per-document concurrency from ConcurrencyConfig.max_threads.
result_format XBERGAlefHandle XBERG_UNIFIED Result structure format Controls whether results are returned in unified format (default) with all content in the content field, or element-based format with semantic elements (for Unstructured-compatible output).
security_limits XBERGAlefHandle NULL Security limits for archive extraction. Controls maximum archive size, compression ratio, file count, and other security thresholds to prevent decompression bomb attacks. Also caps nesting depth, iteration count, entity / token length, total content size, decoded image allocation, and table cell count for every extraction path that ingests user-controlled bytes. When NULL, default limits are used.
max_embedded_file_bytes uint64_t* 52428800 Maximum uncompressed size in bytes for a single embedded file before recursive extraction is attempted (default: 50 MiB). Applies to embedded objects inside OOXML containers (DOCX, PPTX) and to email attachments processed via recursive extraction. Files that exceed this limit are skipped with a ProcessingWarning rather than passed to the extraction pipeline, preventing a single oversized embedded object from consuming unbounded memory or time. Set to NULL to disable the per-embedded-file cap (falls back to security_limits.max_archive_size as the only guard).
output_format XBERGAlefHandle XBERG_PLAIN Content text format (default: Plain). Controls the format of the extracted content: - Plain: Raw extracted text (default) - Markdown: Markdown formatted output - Djot: Djot markup format (requires djot feature) - Html: HTML formatted output When set to a structured format, extraction results will include formatted output. The formatted_content field may be populated when format conversion is applied.
escape_markdown int32_t true Escape Markdown special characters in rendered prose (default: true). When output_format is Markdown or Djot, the renderer backslash-escapes CommonMark-significant leading characters (e.g. -, #) so that literal text such as #06-18 or - clause round-trips safely through a CommonMark parser instead of being reinterpreted as a heading or list marker. Table cell text is never escaped, so escaped prose can look inconsistent with table cells containing the same characters. Set this to false to disable prose escaping and make content, pages[].content, and chunks[].content read identically to table cell text — useful for LLM prompts or search indexing where CommonMark round-tripping does not matter. Defaults to true to preserve existing behavior.
table_anchors int32_t false Emit an opt-in anchor marker before each table’s rendered Markdown block (default: false). When output_format is Markdown (or Djot) and this is true, the renderer inserts a [TABLE:{table_id}] marker immediately before each table’s Markdown in content, pages[].content, and chunks[].content, where table_id matches the corresponding entry’s table_id. This lets a consumer reconcile a rendered Markdown table block with its structured tables[] entry. Defaults to false so existing output is byte-identical unless explicitly enabled.
jupyter_cell_rendering XBERGAlefHandle XBERG_BOTH Controls how Jupyter notebook (.ipynb) code cells are rendered. - Both (default): code source plus the notebook’s saved outputs - Source: only the code source (fenced code blocks) - Outputs: only the saved outputs Cells are never executed; Outputs/Both surface only outputs already stored in the notebook.
apply_notebook_cell_tags int32_t true Apply Jupyter Book/MyST cell visibility tags while rendering notebooks. When enabled, remove-cell/hide-cell, remove-input/hide-input, and remove-output/hide-output suppress the corresponding saved source or output. Cells are never executed, and their metadata remains available even when their rendered content is suppressed. Defaults to true. Set this to false to preserve all saved notebook content regardless of cell tags.
layout XBERGAlefHandle NULL Layout detection configuration (None = layout detection disabled). When set, PDF pages and images are analyzed for document structure (headings, code, formulas, tables, figures, etc.) using RT-DETR models via ONNX Runtime. For PDFs, layout hints override paragraph classification in the markdown pipeline. For images, per-region OCR is performed with markdown formatting based on detected layout classes. Requires the layout-detection feature to run inference; the field is present whenever the layout-types feature is active (which includes layout-detection as well as the no-ORT target groups).
transcription XBERGAlefHandle NULL Transcription (speech-to-text) configuration for audio/video files. When set and enabled, files with audio/video MIME types (mp3, mp4, m4a, wav, webm, etc.) are routed to the Whisper-based transcription pipeline. The actual heavy dependencies are only active under the transcription feature; the field is visible under transcription-types (including on WASM and Android targets that use the no-ORT preset). Default: NULL (transcription disabled). This is an additive, non-breaking change.
use_layout_for_markdown int32_t false Run layout detection on the non-OCR PDF markdown path. When true and layout is Some(_), layout regions inform reading order, region grouping, and table detection while native font/tag semantics remain authoritative for headings, lists, code, and formulas. OCR layout classification is unchanged. This improves structural output at the cost of inference latency (~150-300ms/page CPU, ~20-50ms/page GPU). Default: false. Requires the layout-detection feature.
include_document_structure int32_t false Enable structured document tree output. When true, populates the document field on ExtractedDocument with a hierarchical DocumentStructure containing heading-driven section nesting, table grids, content layer classification, and inline annotations. Independent of result_format — can be combined with Unified or ElementBased.
acceleration XBERGAlefHandle NULL Hardware acceleration configuration for ONNX Runtime models. Controls execution provider selection for layout detection and embedding models. When NULL, uses platform defaults (CoreML on macOS, CUDA on Linux, CPU on Windows).
cache_namespace const char* NULL Cache namespace for tenant isolation. When set, cache entries are stored under {cache_dir}/{namespace}/. Must be alphanumeric, hyphens, or underscores only (max 64 chars). Different namespaces have isolated cache spaces on the same filesystem.
cache_ttl_secs uint64_t* NULL Per-request cache TTL in seconds. Overrides the global max_age_days for this specific extraction. When 0, caching is completely skipped (no read or write). When NULL, the global TTL applies.
email XBERGAlefHandle NULL Email extraction configuration (None = use defaults). Currently supports configuring the fallback codepage for MSG files that do not specify one. See EmailConfig for details.
csv XBERGAlefHandle NULL CSV/TSV extraction configuration (None = use defaults). Lets callers set an explicit delimiter and declare comment-line prefixes to skip, instead of relying solely on delimiter auto-detection. See CsvConfig for details.
geojson XBERGAlefHandle NULL GeoJSON extraction configuration (None = bounded summary). By default, GeoJSON coordinates are replaced by aggregate counts and bounds so large geometry arrays do not become unbounded rendered output. Set include_full_coordinates explicitly to retain the legacy full-coordinate output.
concurrency XBERGAlefHandle NULL Concurrency limits for constrained environments (None = use defaults). Controls Rayon thread pool size, ONNX Runtime intra-op threads, and the combined document/inner-task budget for batch extraction. See ConcurrencyConfig for details.
url XBERGAlefHandle — URL ingestion and crawl configuration.
max_archive_depth uintptr_t 3 Maximum recursion depth for archive extraction (default: 3). Set to 0 to disable recursive extraction (legacy behavior).
tree_sitter XBERGAlefHandle NULL Tree-sitter language pack configuration (None = tree-sitter disabled). When set, enables code file extraction using tree-sitter parsers. Controls grammar download behavior and code analysis options.
structured_extraction XBERGAlefHandle NULL Structured extraction via LLM (None = disabled). When set, the extracted document content is sent to an LLM with the provided JSON schema. The structured response is stored in ExtractedDocument.structured_output.
ner XBERGAlefHandle NULL Named-entity recognition configuration. When set, the NER post-processor runs at the Middle stage and populates ExtractedDocument.entities.
redaction XBERGAlefHandle NULL Redaction / anonymisation configuration. When set, the redaction post-processor runs at the Late stage and rewrites every textual field in ExtractedDocument, emitting an audit trail in ExtractedDocument.redaction_report.
summarization XBERGAlefHandle NULL Summarisation configuration. When set, the summarisation post-processor runs at the Middle stage and populates ExtractedDocument.summary.
translation XBERGAlefHandle NULL Translation configuration. When set, the translation post-processor runs at the Middle stage and populates ExtractedDocument.translation.
page_classification XBERGAlefHandle NULL Per-page classification configuration. When set, the classification post-processor runs at the Middle stage and populates ExtractedDocument.page_classifications.
chunk_classification XBERGAlefHandle NULL Per-chunk multi-label classification configuration. When set, the chunk-classification post-processor runs at the Middle stage (after chunking) and populates ChunkMetadata.classifications on every chunk.
captioning XBERGAlefHandle NULL VLM captioning configuration for extracted images. When set, the captioning post-processor runs at the Middle stage and writes a caption into each ExtractedImage.caption.
qr_codes int32_t* NULL Enable QR-code detection in extracted images. When true, the QR post-processor runs at the Middle stage and populates ExtractedImage.qr_codes.

C representation: XBERGExtractionDiff is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractionDiff does not appear anywhere in the generated header.

The complete diff between two ExtractedDocument values.

Field Type Default Description
content_diff const char* NULL Unified-diff hunks for the content field. Empty when the content is identical.
tables_added const char* NULL Tables present in b but not in a (by index position, excess right-side tables).
tables_removed const char* NULL Tables present in a but not in b (by index position, excess left-side tables).
tables_changed const char* NULL Cell-level changes for table pairs that share the same index and dimensions.
metadata_changed const char* — Metadata difference, encoded as a JSON object with three top-level keys: added (keys present in b but not a), removed (keys present in a but not b), and changed (keys whose values differ — each entry is { "from": <value-in-a>, "to": <value-in-b> }). This is NOT RFC 6902 JSON Patch — we deliberately chose a flatter shape to avoid pulling in a json-patch crate. If you need RFC 6902 semantics (with JSON Pointer paths) feed a.metadata and b.metadata to your preferred json-patch impl directly.
embedded_changes XBERGAlefHandle — Changes to embedded archive children.

C representation: XBERGExtractionErrorItem is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractionErrorItem does not appear anywhere in the generated header.

Non-fatal per-input extraction error captured by ExtractionResult.

Field Type Default Description
index uintptr_t — Input index in the original request.
code uint32_t — Stable numeric error code.
error_type const char* — Stable snake_case error kind.
source const char* — Best-effort source identifier.
message const char* — Error message.

C representation: XBERGExtractionResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractionResult does not appear anywhere in the generated header.

Unified extraction result envelope.

Field Type Default Description
results const char* NULL Extracted documents in discovery order.
errors const char* NULL Non-fatal per-input errors.
summary XBERGAlefHandle — Aggregate counts for the operation.
crawl_final_urls const char* NULL Final URLs reached after redirects during URL ingestion.
crawl_redirect_count uintptr_t — Total redirects followed while fetching or crawling URLs.
crawl_unique_normalized_urls const char* NULL Unique normalized URLs discovered by crawls.

C representation: XBERGExtractionSummary is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGExtractionSummary does not appear anywhere in the generated header.

Summary for a unified extraction call.

Field Type Default Description
inputs uintptr_t — Number of inputs submitted by the caller.
results uintptr_t — Number of extraction results produced.
errors uintptr_t — Number of per-input errors.
remote_urls uintptr_t — Number of URI inputs that resolved to remote HTTP(S) URLs.
pages_crawled uintptr_t — Number of HTML pages crawled or scraped.
documents_downloaded uintptr_t — Number of downloaded non-HTML documents extracted from URLs.

C representation: XBERGFictionBookMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGFictionBookMetadata does not appear anywhere in the generated header.

FictionBook (FB2) metadata.

Field Type Default Description
genres const char* NULL Genre tags as declared in the FB2 <genre> elements.
sequences const char* NULL Book series (sequence) names, if any.
annotation const char* NULL Short annotation / summary from the FB2 <annotation> element.

C representation: XBERGFileExtractionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGFileExtractionConfig does not appear anywhere in the generated header.

Per-file extraction configuration overrides for batch processing.

All fields are Option<T> — NULL means “use the batch-level default.” This type is used by config and extract_batch to allow heterogeneous extraction settings within a single batch.

The following ExtractionConfig fields are batch-level only and cannot be overridden per file:

  • max_concurrent_extractions — controls batch parallelism
  • use_cache — global caching policy
  • acceleration — shared ONNX execution provider
  • security_limits — global archive security policy
Field Type Default Description
mime_detection_policy XBERGAlefHandle NULL Override MIME inference policy for this file.
enable_quality_processing int32_t* NULL Override quality post-processing for this file.
ocr XBERGAlefHandle NULL Override OCR configuration for this file (None in the Option = use batch default).
force_ocr int32_t* NULL Override force OCR for this file.
ocr_strategy XBERGAlefHandle NULL Override the OCR page-selection strategy for this file.
force_ocr_pages const char* NULL Override force OCR pages for this file (1-indexed page numbers).
disable_ocr int32_t* NULL Override disable OCR for this file.
chunking XBERGAlefHandle NULL Override chunking configuration for this file.
content_filter XBERGAlefHandle NULL Override content filtering configuration for this file.
images XBERGAlefHandle NULL Override image extraction configuration for this file.
pdf_options XBERGAlefHandle NULL Override PDF options for this file.
token_reduction XBERGAlefHandle NULL Override token reduction for this file.
language_detection XBERGAlefHandle NULL Override language detection for this file.
pages XBERGAlefHandle NULL Override page extraction for this file.
keywords XBERGAlefHandle NULL Override keyword extraction for this file.
postprocessor XBERGAlefHandle NULL Override post-processor for this file.
html_options XBERGAlefHandle NULL Override HTML conversion options for this file.
html_output XBERGAlefHandle NULL Override styled HTML output configuration for this file.
result_format XBERGAlefHandle NULL Override result format for this file.
output_format XBERGAlefHandle NULL Override output content format for this file.
include_document_structure int32_t* NULL Override document structure output for this file.
layout XBERGAlefHandle NULL Override layout detection for this file.
transcription XBERGAlefHandle NULL Transcription configuration (see ExtractionConfig for docs).
timeout_secs uint64_t* NULL Override per-file extraction timeout in seconds. When set, the extraction for this file will be canceled after the specified duration. A timed-out file produces an error result without affecting other files in the batch.
tree_sitter XBERGAlefHandle NULL Override tree-sitter configuration for this file.
structured_extraction XBERGAlefHandle NULL Override structured extraction configuration for this file. When set, enables LLM-based structured extraction with a JSON schema for this specific file. The extracted content is sent to a VLM/LLM and the response is parsed according to the provided schema.
url XBERGAlefHandle NULL Override URL ingestion and crawl configuration for this file.
ner XBERGAlefHandle NULL Override named-entity recognition configuration for this file.
redaction XBERGAlefHandle NULL Override redaction configuration for this file.
summarization XBERGAlefHandle NULL Override summarization configuration for this file.
translation XBERGAlefHandle NULL Override translation configuration for this file.
page_classification XBERGAlefHandle NULL Override per-page classification configuration for this file.
chunk_classification XBERGAlefHandle NULL Override per-chunk classification configuration for this file.
captioning XBERGAlefHandle NULL Override VLM captioning configuration for this file.
qr_codes int32_t* NULL Override QR-code detection for this file.

C representation: XBERGFootnote is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGFootnote does not appear anywhere in the generated header.

Footnote in Djot.

Field Type Default Description
label const char* — Footnote label
content const char* — Footnote content blocks

C representation: XBERGFootnoteAnchor is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGFootnoteAnchor does not appear anywhere in the generated header.

A footnote anchor reference in markdown text.

Represents a [^label] use-site (not a definition).

Field Type Default Description
label const char* — The label of the footnote reference (e.g., “1” in [^1]).
offset uintptr_t — Byte offset of the anchor in the markdown text.

C representation: XBERGFootnoteConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGFootnoteConfig does not appear anywhere in the generated header.

Configuration for markdown footnote and citation parsing.

Field Type Default Description
parse_citations int32_t true Whether to parse the structured citation block (default: true). When enabled, the parser will look for and extract citations from the block after --- + <!-- citations ... -->.

C representation: XBERGFootnoteDefinition is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGFootnoteDefinition does not appear anywhere in the generated header.

A footnote definition from markdown text.

Represents [^label]: content declarations (including multi-line continuations).

Field Type Default Description
label const char* — The label of the footnote (e.g., “1” in [^1]: ...).
content const char* — The full content of the footnote definition.
offset uintptr_t — Byte offset of the definition line in the markdown text.

C representation: XBERGFormattedBlock is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGFormattedBlock does not appear anywhere in the generated header.

Block-level element in a Djot document.

Represents structural elements like headings, paragraphs, lists, code blocks, etc.

Field Type Default Description
block_type XBERGAlefHandle — Type of block element
level uintptr_t* NULL Heading level (1-6) for headings, or nesting level for lists
inline_content const char* — Inline content within the block
attributes XBERGAlefHandle NULL Element attributes (classes, IDs, key-value pairs)
language const char* NULL Language identifier for code blocks
code const char* NULL Raw code content for code blocks
children const char* /* serde(default) */ Nested blocks for containers (blockquotes, list items, divs)

C representation: XBERGFormula is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGFormula does not appear anywhere in the generated header.

A mathematical formula extracted from a document.

Three kinds of sources populate this type. Layout-guided OCR detects formula regions and recognizes them; those formulas carry a bbox and a page. VLM OCR recognizes formulas in transcribed text without layout, so its formulas carry no geometry. Markup extraction (DOCX, PPTX, ODT, EPUB, HTML, JATS, LaTeX, Markdown, and related formats) converts embedded math to LaTeX, also without geometry.

Field Type Default Description
latex const char* — LaTeX source of the formula, without surrounding $$ delimiters. Markup converters and formula OCR produce real LaTeX. The native PDF layout path stores the plain text of a detected formula region, which keeps the original Unicode math characters instead of LaTeX commands. To render the formula in Markdown or other formats, wrap it in $$..$$.
bbox XBERGAlefHandle /* serde(default) */ Bounding box of the formula region on its page. NULL for markup sources. PDF OCR sources report PDF point coordinates with the origin at the bottom-left of the page, comparable to native PDF geometry. Image sources, and PDF pages whose geometry is unavailable, report pixels of the image the OCR backend saw. The C FFI reports an absent bbox as a null pointer.
page uint32_t* /* serde(default) */ 1-indexed page number the formula appears on. NULL when the source format has no page concept. The C FFI reports an absent page as 0.

Since: v1.1

C representation: XBERGGeoJsonExtractionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGGeoJsonExtractionConfig does not appear anywhere in the generated header.

Configuration for GeoJSON extraction.

GeoJSON coordinates can dominate extraction output and duplicate large geometry payloads in rendered content and metadata. The default therefore emits a bounded aggregate summary. Set Self.include_full_coordinates only when callers need every coordinate in the extracted text and accept output proportional to the input.

Field Type Default Description
include_full_coordinates int32_t — Include every coordinate in rendered content and flattened_fields metadata. Defaults to false. When false, extraction reports feature, property, geometry, position-count, bounds, truncation, and discarded-category metadata and emits a ProcessingWarning for every GeoJSON input.

C representation: XBERGGridCell is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGGridCell does not appear anywhere in the generated header.

Individual grid cell with position and span metadata.

Field Type Default Description
content const char* — Cell text content.
row uint32_t — Zero-indexed row position.
col uint32_t — Zero-indexed column position.
row_span uint32_t 1 Number of rows this cell spans.
col_span uint32_t 1 Number of columns this cell spans.
is_header int32_t /* serde(default) */ Whether this is a header cell.
bbox XBERGAlefHandle NULL Bounding box for this cell (if available).
heading_level uint8_t* /* serde(default) */ Outline level (1-6) of the heading style this cell’s text carries, when it has one. A DOCX banner row – row 0, one cell spanning the grid, styled Heading1..Heading6 – is what Word’s navigation pane and a TOC field treat as the document’s outline, but as a table cell it reached consumers as anonymous text (GH#1587). content is deliberately left as the bare cell text: prefixing it with # would put a markdown heading inside a table cell, which is invalid where it lands and changes text every existing consumer already reads. This field is the signal instead, so a caller can decide for itself whether a heading 2 in a banner row is a section title or a column label.
style_name const char* /* serde(default) */ Human-readable name of the paragraph style applied to this cell’s text (heading 2). Carries the style even when it resolves to no outline level, so a caller can key on a named style this library does not map to a heading. See GridCell.heading_level.

C representation: XBERGHeaderMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHeaderMetadata does not appear anywhere in the generated header.

Header/heading element metadata.

Field Type Default Description
level uint8_t — Header level: 1 (h1) through 6 (h6)
text const char* — Normalized text content of the header
id const char* NULL HTML id attribute if present
depth uint32_t — Document tree depth at the header element
html_offset uint32_t — Byte offset in original HTML document

C representation: XBERGHeadingContext is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHeadingContext does not appear anywhere in the generated header.

Heading context for a chunk within a Markdown document.

Contains the heading hierarchy from document root to this chunk’s section.

Field Type Default Description
headings const char* — The heading hierarchy from document root to this chunk’s section. Index 0 is the outermost (h1), last element is the most specific.

C representation: XBERGHeadingLevel is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHeadingLevel does not appear anywhere in the generated header.

A single heading in the hierarchy.

Field Type Default Description
level uint8_t — Heading depth (1 = h1, 2 = h2, etc.)
text const char* — The text content of the heading.

C representation: XBERGHeuristicsConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHeuristicsConfig does not appear anywhere in the generated header.

Configuration for document chunking and analysis heuristics.

Every threshold is a public field so callers can override any subset via struct-update syntax: HeuristicsConfig { text_layer_threshold: 0.5, ..the default constructor }.

Field Type Default Description
enable_pdf_text_heuristics int32_t true Enable PDF text-layer detection heuristics. When true, PDFs with a substantial text layer will skip chunking. Default: true.
text_layer_threshold float 0.7 Minimum fraction of pages that must have text to skip chunking. Range 0.0..=1.0. Default: 0.7 (70 % of pages).
file_size_threshold_bytes uint64_t 10485760 File size threshold in bytes for considering chunking. Files smaller than this are processed without chunking. Default: 10 MiB (10 × 1 024 × 1 024).
page_count_threshold uint32_t 50 Page count threshold for considering chunking. Documents with fewer pages are processed without chunking. Default: 50.
target_pages_per_chunk uint32_t 10 Target number of pages per chunk for optimal parallel processing. Default: 10.
max_pages_per_chunk uint32_t 25 Hard cap on pages per chunk. No chunk will exceed this limit. Must be ≥ target_pages_per_chunk. Default: 25.
disk_processing_threshold_bytes uint64_t 52428800 File size threshold for disk-based processing. Files larger than this are buffered to disk to prevent OOM. Default: 50 MiB (50 × 1 024 × 1 024).
min_chars_per_page uint32_t 50 Minimum characters per page to consider a page as having text. Default: 50.
max_xlsx_sheet_count uint32_t 200 Maximum sheet count allowed in an XLSX workbook. Workbooks beyond this are rejected pre-extraction to avoid OOM / abusive billing inflation. Default: 200.
max_xlsx_workbook_cells uint64_t 5000000 Maximum cell count (sheets × rows × columns approximation) in an XLSX workbook. Default: 5 000 000 (≈ 200 sheets × 25 k cells).
max_pptx_embedded_count uint32_t 50 Maximum number of OLE-embedded objects extractable from a single PPTX or DOCX. Protects against zip-bomb-style nested-document abuse. Default: 50.

C representation: XBERGHierarchicalBlock is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHierarchicalBlock does not appear anywhere in the generated header.

Field Type Default Description
text const char* — The text content of this block
font_size float — The font size of the text in this block
level const char* — The hierarchy level of this block (H1-H6 or Body) Levels correspond to HTML heading tags: - “h1”: Top-level heading - “h2”: Secondary heading - “h3”: Tertiary heading - “h4”: Quaternary heading - “h5”: Quinary heading - “h6”: Senary heading - “body”: Body text (no heading level)
bbox XBERGAlefHandle NULL Bounding box information for the block Contains left, top, right, and bottom coordinates in PDF units.

C representation: XBERGHierarchicalBoundingBox is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHierarchicalBoundingBox does not appear anywhere in the generated header.

A text block with hierarchy level assignment.

Represents a block of text with semantic heading information extracted from font size clustering and hierarchical analysis.

Field Type Default Description
left float — Left coordinate.
top float — Top coordinate.
right float — Right coordinate.
bottom float — Bottom coordinate.

C representation: XBERGHierarchyConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHierarchyConfig does not appear anywhere in the generated header.

Hierarchy extraction configuration for PDF text structure analysis.

Enables extraction of document hierarchy levels (H1-H6) based on font size clustering and semantic analysis. When enabled, hierarchical blocks are included in page content.

Field Type Default Description
enabled int32_t true Enable hierarchy extraction
k_clusters uintptr_t 3 Number of font size clusters to use for hierarchy levels (1-7) Default: 3, which provides two heading levels plus body text. Larger values create more fine-grained hierarchy levels.
include_bbox int32_t true Include bounding box information in hierarchy blocks

C representation: XBERGHtmlMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHtmlMetadata does not appear anywhere in the generated header.

HTML metadata extracted from HTML documents.

Includes document-level metadata, Open Graph data, Twitter Card metadata, and extracted structural elements (headers, links, images, structured data).

Field Type Default Description
title const char* NULL Document title from <title> tag
description const char* NULL Document description from <meta name="description"> tag
keywords const char* NULL Document keywords from <meta name="keywords"> tag, split on commas
author const char* NULL Document author from <meta name="author"> tag
canonical_url const char* NULL Canonical URL from <link rel="canonical"> tag
base_href const char* NULL Base URL from <base href=""> tag for resolving relative URLs
language const char* NULL Document language from lang attribute
text_direction XBERGAlefHandle NULL Document text direction from dir attribute
open_graph const char* NULL Open Graph metadata (og:* properties) for social media Keys like “title”, “description”, “image”, “url”, etc.
twitter_card const char* NULL Twitter Card metadata (twitter:* properties) Keys like “card”, “site”, “creator”, “title”, “description”, “image”, etc.
meta_tags const char* NULL Additional meta tags not covered by specific fields Keys are meta name/property attributes, values are content
headers const char* NULL Extracted header elements with hierarchy
links const char* NULL Extracted hyperlinks with type classification
images const char* NULL Extracted images with source and dimensions
structured_data const char* NULL Extracted structured data blocks

C representation: XBERGHtmlOutputConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGHtmlOutputConfig does not appear anywhere in the generated header.

Configuration for styled HTML output.

When set on html_output alongside output_format = OutputFormat.Html, the pipeline builds a StyledHtmlRenderer instead of the plain comrak-based renderer.

Field Type Default Description
css const char* NULL Inline CSS string injected into the output after the theme stylesheet. Concatenated after css_file content when both are set.
css_file const char* NULL Path to a CSS file loaded once at renderer construction time. Concatenated before css when both are set.
theme XBERGAlefHandle XBERG_UNSTYLED Built-in colour/typography theme. Default: HtmlTheme.Unstyled.
class_prefix const char* "kb-" CSS class prefix applied to every emitted class name. Default: "kb-". Change this if your host application already uses classes that start with kb-.
embed_css int32_t true When true (default), write the resolved CSS into a <style> block immediately after the opening <div class="{prefix}doc">. Set to false to emit only the structural markup and wire up your own stylesheet targeting the kb-* class names.

C representation: XBERGImageDimensions is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGImageDimensions does not appear anywhere in the generated header.

Image dimensions in pixels.

Field Type Default Description
width uint32_t — Width in pixels.
height uint32_t — Height in pixels.

C representation: XBERGImageDpi is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGImageDpi does not appear anywhere in the generated header.

Horizontal and vertical image resolution in dots per inch.

Field Type Default Description
horizontal double — Horizontal resolution.
vertical double — Vertical resolution.

C representation: XBERGImageExtractionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGImageExtractionConfig does not appear anywhere in the generated header.

Image extraction configuration.

Field Type Default Description
extract_images int32_t true Extract images from documents
target_dpi int32_t 300 Target DPI for image normalization
max_image_dimension int32_t 4096 Maximum dimension for images (width or height)
inject_placeholders int32_t true Whether to inject image reference placeholders into markdown output. When true (default), image references like !Image 1 are appended to the markdown. Set to false to extract images as data without polluting the markdown output.
auto_adjust_dpi int32_t true Automatically adjust DPI based on image content
min_dpi int32_t 72 Minimum DPI threshold
max_dpi int32_t 600 Maximum DPI threshold
max_images_per_page uint32_t* NULL Maximum number of image objects to extract per PDF page. Some PDFs (e.g. technical diagrams stored as thousands of raster fragments) can trigger extremely long or indefinite extraction times when every image object on a dense page is decoded individually via the PDF extractor. Setting this limit causes xberg to stop collecting individual images once the count per page reaches the cap and emit a warning instead. NULL (default) means no limit — all images are extracted.
classify int32_t false When true, extracted images are classified by kind and grouped into clusters where they appear to belong to one figure. Defaults to false — opt in explicitly to avoid unexpected ML overhead.
include_page_rasters int32_t false When true, full-page renders produced during OCR preprocessing are captured and returned as ImageKind.PageRaster entries in ExtractedDocument.images. PDF + OCR only. No rasters are captured for non-PDF inputs or when the document-level OCR bypass is active (whole-document backend). When OCR is enabled and this flag is set but the active backend skips per-page rendering, a ProcessingWarning is emitted in ExtractedDocument.processing_warnings. Defaults to false. Enable when downstream consumers need page thumbnails (e.g. citation previews, visual grounding).
run_ocr_on_images int32_t true Run OCR on extracted images and include the recognized text in the document content. When true (default) and ExtractionConfig.ocr is configured, extracted images are processed with the configured OCR backend. Set to false to extract images without OCR processing, even when OCR is enabled.
ocr_text_only int32_t false When true, image OCR results are rendered as plain text without the ![...](...) markdown placeholder. Only takes effect when run_ocr_on_images is also true.
append_ocr_text int32_t false When true and ocr_text_only is false, append the OCR text after the image placeholder in the rendered output.
output_format XBERGAlefHandle XBERG_NATIVE Target format for re-encoding extracted images. When set to anything other than Native, each extracted image is re-encoded to the requested format before being returned. This lets callers receive uniform output without duplicating encode logic downstream. Defaults to Native — no re-encode pass is performed and ExtractedImage.format reflects the source extractor’s output.
svg XBERGAlefHandle — SVG-specific knobs for the image-encode pipeline. Controls sanitization and rasterization DPI when the source or output format is SVG. Only available when the svg feature is active.
include_data_base64 int32_t false When true, populate ExtractedImage.data_base64 with a Base64-encoded copy of the raw image bytes. Useful for JSON-only clients that cannot efficiently parse the default integer-array serialization of data. Defaults to false; enabling it doubles the in-memory image representation for the duration of the response.

C representation: XBERGImageMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGImageMetadata does not appear anywhere in the generated header.

Image metadata extracted from image files.

Includes dimensions, format, and EXIF data.

Field Type Default Description
width uint32_t — Image width in pixels
height uint32_t — Image height in pixels
format const char* — Image format (e.g., “PNG”, “JPEG”, “TIFF”)
exif const char* NULL EXIF metadata tags

C representation: XBERGImageMetadataType is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGImageMetadataType does not appear anywhere in the generated header.

Image element metadata.

Field Type Default Description
src const char* — Image source (URL, data URI, or SVG content)
alt const char* NULL Alternative text from alt attribute
title const char* NULL Title attribute
dimensions XBERGAlefHandle NULL Image dimensions if available.
image_type XBERGAlefHandle — Image type classification
attributes const char* — Additional attributes as key-value pairs.

C representation: XBERGImagePreprocessingConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGImagePreprocessingConfig does not appear anywhere in the generated header.

Image preprocessing configuration for OCR.

These settings control how images are preprocessed before OCR to improve text recognition quality. Different preprocessing strategies work better for different document types.

Field Type Default Description
target_dpi int32_t 300 Target DPI for the image (300 is standard, 600 for small text). For a PDF page, this resamples the already-rendered raster; it does not make the page re-render at a higher native resolution. A value above what image.preprocessing.calculate_target_dpi’s memory/dimension clamp allows is silently capped (GH#1786: 400 and 600 both clamped to the same ~372 on a Letter page and produced byte-identical output), so upscaling interpolated pixels this way adds no detail. To change the actual PDF render resolution, set images.target_dpi on ExtractionConfig instead (image.dpi.effective_pdf_render_dpi).
auto_rotate int32_t false Auto-detect and correct image rotation.
deskew int32_t true Correct skew (tilted images). Must be false with none or off binarization.
denoise int32_t false Remove noise from the image.
contrast_enhance int32_t false Enhance contrast for better text visibility. With none or off binarization, this applies background normalization and sharpening while preserving grayscale pixels.
binarization_method const char* "otsu" Binarization method: “none” (alias “off”), “otsu”, “sauvola”, or “adaptive”. deskew must be false when this is none or off.
invert_colors int32_t false Invert colors (white text on black → black on white).
normalize_shaded_rows int32_t false Normalize shaded table rows (e.g. a subtotal row on a light or dark fill) before binarization, so each shaded band is stretched to its own dark-text-on-white polarity instead of being lost to a single whole-page threshold (GH#1785). This is a per-band step, not a replacement for binarization_method: no single whole-page method recovers every fill color, and the per-band step itself can regress a row style it does not fully model (e.g. a mid-grey fill with white text), so it defaults to false rather than being enabled unconditionally. Has no effect when binarization_method is "none" or "off" and contrast_enhance is true: that path re-reads the original image to apply background normalization and sharpening, discarding the per-band result. A WARN is emitted when the option is requested in that combination (GH#1837).

C representation: XBERGImagePreprocessingMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGImagePreprocessingMetadata does not appear anywhere in the generated header.

Field Type Default Description
original_dimensions XBERGAlefHandle — Original image dimensions in pixels.
original_dpi XBERGAlefHandle — Original image resolution.
target_dpi int32_t — Target DPI from configuration
scale_factor double — Scaling factor applied to the image
auto_adjusted int32_t — Whether DPI was auto-adjusted based on content
final_dpi int32_t — Final DPI after processing
new_dimensions XBERGAlefHandle NULL New dimensions after resizing (if resized).
resample_method const char* — Resampling algorithm used (“LANCZOS3”, “CATMULLROM”, etc.)
dimension_clamped int32_t — Whether dimensions were clamped to max_image_dimension
calculated_dpi int32_t* NULL Calculated optimal DPI (if auto_adjust_dpi enabled)
skipped_resize int32_t — Whether resize was skipped (dimensions already optimal)
resize_error const char* NULL Error message if resize failed

C representation: XBERGInlineElement is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGInlineElement does not appear anywhere in the generated header.

Inline element within a block.

Represents text with formatting, links, images, etc.

Field Type Default Description
element_type XBERGAlefHandle — Type of inline element
content const char* — Text content
attributes XBERGAlefHandle NULL Element attributes
metadata const char* NULL Additional metadata (e.g., href for links, src/alt for images)

C representation: XBERGJatsMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGJatsMetadata does not appear anywhere in the generated header.

JATS (Journal Article Tag Suite) metadata.

Field Type Default Description
copyright const char* NULL Copyright statement from the article’s <permissions> element.
license const char* NULL Open-access license URI from the article’s <license> element.
history_dates const char* NULL Publication history dates keyed by event type (e.g. "received", "accepted").
contributor_roles const char* NULL Authors and contributors with their stated roles.

C representation: XBERGKeyValueAttribute is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGKeyValueAttribute does not appear anywhere in the generated header.

A string key-value attribute.

Field Type Default Description
key const char* — Attribute name.
value const char* — Attribute value.

C representation: XBERGKeyword is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGKeyword does not appear anywhere in the generated header.

Extracted keyword with metadata.

Field Type Default Description
text const char* — The keyword text.
score float — Relevance score (higher is better, algorithm-specific range).
algorithm XBERGAlefHandle — Algorithm that extracted this keyword.
positions const char* NULL Optional positions where keyword appears in text (character offsets).

C representation: XBERGKeywordConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGKeywordConfig does not appear anywhere in the generated header.

Keyword extraction configuration.

Field Type Default Description
algorithm XBERGAlefHandle XBERG_YAKE Algorithm to use for extraction.
max_keywords uintptr_t 10 Maximum number of keywords to extract (default: 10).
min_score float 0 Minimum score threshold (0.0-1.0, default: 0.0). Keywords with scores below this threshold are filtered out. Note: Score ranges differ between algorithms.
ngram_range XBERGAlefHandle — N-gram range for keyword extraction (min, max). (1, 1) = unigrams only (1, 2) = unigrams and bigrams (1, 3) = unigrams, bigrams, and trigrams (default)
language const char* "en" Language code for stopword filtering (e.g., “en”, “de”, “fr”). If None, no stopword filtering is applied.
yake_params XBERGAlefHandle NULL YAKE-specific tuning parameters.
rake_params XBERGAlefHandle NULL RAKE-specific tuning parameters.

Since: v1.1

C representation: XBERGLanguageConfidence is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLanguageConfidence does not appear anywhere in the generated header.

Structured per-language detection result: confidence, document share, and script — the information the ISO-code-only detected_languages list cannot convey (#261).

Populated by the language-detection processor alongside detected_languages, with one entry per language, in the same order as detected_languages.

Field Type Default Description
language const char* — ISO 639-3 language code, matching the corresponding entry in detected_languages.
confidence double — Confidence for this language, in [0.0, 1.0]. In single-language mode this is whatlang’s Info.confidence() for the whole document. In multi-language mode this is the average whatlang confidence across the document’s 200-character chunks that were classified as this language.
proportion double — Share of the document’s analyzed content classified as this language, in [0.0, 1.0]. In single-language mode this is always 1.0. In multi-language mode this is the fraction of 200-character chunks classified as this language (chunks that did not meet min_confidence for any language are excluded from the count but still count toward the denominator).
script const char* — Writing system whatlang detected for this language (e.g. "Latin", "Cyrillic").
reliable int32_t — Whether this detection is considered reliable. In single-language mode this is whatlang’s own Info.is_reliable() (confidence above whatlang’s internal 0.9 threshold). In multi-language mode this is the chunk-averaged confidence above that same 0.9 threshold, since whatlang’s is_reliable() only applies to a single detection.

C representation: XBERGLanguageDetectionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLanguageDetectionConfig does not appear anywhere in the generated header.

Language detection configuration.

Field Type Default Description
enabled int32_t true Enable language detection
min_confidence double 0.8 Minimum confidence threshold (0.0-1.0)
detect_multiple int32_t false Detect multiple languages in the document

C representation: XBERGLateInteractionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLateInteractionConfig does not appear anywhere in the generated header.

Configuration for the late-interaction (ColBERT) pipeline.

Controls which model to use, batching, and download/cache behavior for the local ONNX ColBERT model.

Field Type Default Description
model XBERGAlefHandle XBERG_PRESET The late-interaction model to use (defaults to the “gte-moderncolbert” preset).
batch_size uintptr_t 16 Batch size for local ONNX inference. ColBERT emits a [seq, dim] multi-vector embedding per document, so memory scales with batch size — keep this modest.
max_length uintptr_t 512 Maximum token sequence length for the tokenizer (documents).
query_max_length uintptr_t 32 Fixed padded length for query augmentation. ColBERT queries are padded (with the mask token, kept attention-live) to exactly this many tokens rather than truncated/left as-is — this is the “query augmentation” trick from the ColBERT paper.
show_download_progress int32_t false Show model download progress (local ONNX path only). When enabled, transfer progress for the model, tokenizer and config files is reported at info level on the xberg.model_download target while they download (#279). A warm Hugging Face cache transfers nothing and so reports nothing. Ignored by LateInteractionModelType.Plugin, which downloads no model.
cache_dir const char* NULL Optional alternate Hugging Face cache root for model files. When unset, hf-hub follows the standard Hugging Face environment and platform cache conventions.
acceleration XBERGAlefHandle NULL Hardware acceleration for the late-interaction ONNX model.
max_embed_duration_secs uint64_t* 60 Maximum wall-clock duration (in seconds) for a single embed call when using LateInteractionModelType.Plugin. NULL disables the timeout.

C representation: XBERGLateInteractionMatch is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLateInteractionMatch does not appear anywhere in the generated header.

A single document match returned by max_sim_rank, with its position in the input and MaxSim score.

Field Type Default Description
index uintptr_t — Position of this document in the original input slice.
score float — MaxSim relevance score. Higher means more relevant to the query.

C representation: XBERGLateInteractionPreset is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLateInteractionPreset does not appear anywhere in the generated header.

Static metadata for a bundled ColBERT preset (WASM/Android-safe, no ORT).

Field Type Default Description
name const char* — Stable preset name referenced from config.
model_repo const char* — HuggingFace repository hosting the ONNX model.
model_file const char* — Path to the ONNX file within the repo.
additional_files const char* — Sibling files that must be downloaded alongside model_file.
max_length uintptr_t — Maximum document token sequence length.
query_max_length uintptr_t — Fixed padded query length (ColBERT query augmentation).
dim uintptr_t — Per-token embedding dimensionality.
description const char* — Human-readable description.

C representation: XBERGLayoutDetection is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLayoutDetection does not appear anywhere in the generated header.

A single layout detection result.

Field Type Default Description
class_name XBERGAlefHandle — Detected layout class (e.g. Table, Text, Title).
confidence float — Detection confidence score in [0.0, 1.0].
bbox XBERGAlefHandle — Bounding box in image pixel coordinates.

C representation: XBERGLayoutDetectionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLayoutDetectionConfig does not appear anywhere in the generated header.

Layout detection configuration.

Controls layout detection behavior in the extraction pipeline. When set on ExtractionConfig, layout detection is enabled for PDF extraction.

Field Type Default Description
strategy XBERGAlefHandle XBERG_ALWAYS Which pages the layout model runs on. Defaults to LayoutStrategy.Always, the historical behavior: every page is rendered and inferred. LayoutStrategy.Auto pre-screens pages with cheap signals and skips the model where it cannot help.
confidence_threshold float* NULL Confidence threshold override (None = use model default).
apply_heuristics int32_t true Whether to apply postprocessing heuristics (default: true).
table_model XBERGAlefHandle XBERG_TATR Table structure recognition model. Controls which model is used for table cell detection within layout-detected table regions. Defaults to TableModel.Tatr.
formula_model XBERGAlefHandle NULL Formula recognition model for layout-detected formula regions. NULL (the default) keeps the plain OCR text of the region. Setting a model converts each formula region crop to LaTeX. Requires the formula-recognition feature; without it the setting is ignored.
table_overlap_preference XBERGAlefHandle XBERG_CONTENT How to resolve overlapping native vs layout tables. When a native table and a layout (TATR/SLANeXT) table overlap on the same region, this controls which one is kept. Defaults to TableOverlapPreference.Content (historical behavior: keep the table with more content). Set to TableOverlapPreference.Native to favor source reading order (higher text F1) over the model’s cell reflow.
acceleration XBERGAlefHandle NULL Hardware acceleration for ONNX models (layout detection + table structure). When set, controls which execution provider (CPU, CUDA, CoreML, TensorRT) is used for inference. Defaults to NULL (auto-select per platform).
enable_chart_understanding int32_t false Route regions classified as charts to the chart-understanding OCR task. When true, layout regions detected as charts are sent to the VLM chart task (data-series/axis recovery) instead of being treated as generic image regions. Defaults to false — chart understanding is opt-in and has no effect on standard text/table extraction scores.

C representation: XBERGLayoutRegion is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLayoutRegion does not appear anywhere in the generated header.

A detected layout region on a page.

When layout detection is enabled, each page may have layout regions identifying different content types (text, pictures, tables, etc.) with confidence scores and spatial positions.

Field Type Default Description
class_name const char* — Layout class name (e.g. “picture”, “table”, “text”, “section_header”).
confidence double — Confidence score from the layout detection model (0.0 to 1.0).
bounding_box XBERGAlefHandle — Bounding box in document coordinate space.
area_fraction double — Fraction of the page area covered by this region (0.0 to 1.0).

C representation: XBERGLinkMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLinkMetadata does not appear anywhere in the generated header.

Link element metadata.

Field Type Default Description
href const char* — The href URL value
text const char* — Link text content (normalized)
title const char* NULL Optional title attribute
link_type XBERGAlefHandle — Link type classification
rel const char* — Rel attribute values
attributes const char* — Additional attributes as key-value pairs.

Since: v1.1

C representation: XBERGLlmBudgetConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLlmBudgetConfig does not appear anywhere in the generated header.

Budget enforcement configuration.

Mirrors liter-llm’s LlmBudgetConfig. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value round-trips through configuration but is not enforced at request time.

Field Type Default Description
global_limit double* NULL Global spend limit in USD.
model_limits const char* NULL Per-model spend limits in USD, keyed by model name.
enforcement const char* NULL Enforcement mode: "hard" (reject over-budget requests) or "soft" (log only).

Since: v1.1

C representation: XBERGLlmCacheConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLlmCacheConfig does not appear anywhere in the generated header.

Response cache configuration.

Mirrors liter-llm’s LlmCacheConfig. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value round-trips through configuration but is not consulted at request time.

Field Type Default Description
max_entries uintptr_t* NULL Maximum number of cached entries.
ttl_seconds uint64_t* NULL Cache entry time-to-live, in seconds.
backend const char* NULL Cache backend name (e.g. "memory", or an opendal scheme).
backend_config const char* NULL Backend-specific configuration key/value pairs.

C representation: XBERGLlmConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLlmConfig does not appear anywhere in the generated header.

Configuration for an LLM provider/model via liter-llm.

Each feature (VLM OCR, VLM embeddings, structured extraction) carries its own LlmConfig, allowing different providers per feature.

Debug is implemented by hand so api_key, header values, and the AWS credentials in BedrockConfig are never printed.

Field Type Default Description
model const char* — Provider/model string using liter-llm routing format. Examples: "openai/gpt-4o", "anthropic/claude-sonnet-4-20250514", "groq/llama-3.1-70b-versatile".
api_key const char* NULL API key for the provider. When NULL, liter-llm falls back to the provider’s standard environment variable (e.g., OPENAI_API_KEY).
base_url const char* NULL Custom base URL override for the provider endpoint.
timeout_secs uint64_t* NULL Request timeout in seconds. When NULL, liter-llm’s built-in 60s default applies, except the VLM OCR path which uses a 300s default (a single page image transcription routinely exceeds 60s). Set explicitly to override.
max_retries uint32_t* NULL Maximum retry attempts (default: 3).
temperature double* NULL Sampling temperature for generation tasks.
max_tokens uint64_t* NULL Maximum tokens to generate.
top_p double* NULL Nucleus sampling parameter for generation tasks, applied to individual requests built from this config. Restricts sampling to the smallest set of tokens whose cumulative probability mass is at least this value; lower is more focused. Validated to [0.0, 1.0] by LlmConfig.validate. Mirrors liter-llm’s ChatCompletionRequest.top_p. A request-time parameter like temperature/max_tokens above, not a client-level setting.
stop const char* NULL Stop sequence(s) that halt token generation, applied to individual requests built from this config. Mirrors liter-llm’s ChatCompletionRequest.stop (types.common.StopSequence), which liter-llm represents as either a single string or a list of strings via an untagged enum. Always expressed here as a list — even one stop sequence is ["..."] — so the field has a single, FFI-friendly shape across every language binding instead of a single-or-list union type. Converted to liter-llm’s StopSequence.Multiple at each request-building call site; see llm.client.to_stop_sequence.
seed int64_t* NULL Random seed for reproducible outputs, applied to individual requests built from this config. Provider support varies — some silently ignore it. Mirrors liter-llm’s ChatCompletionRequest.seed.
presence_penalty double* NULL Presence penalty for generation tasks, applied to individual requests built from this config. Positive values discourage the model from repeating topics already present in the conversation. Validated to [-2.0, 2.0] by LlmConfig.validate. Mirrors liter-llm’s ChatCompletionRequest.presence_penalty.
frequency_penalty double* NULL Frequency penalty for generation tasks, applied to individual requests built from this config. Positive values discourage the model from repeating the same tokens verbatim. Validated to [-2.0, 2.0] by LlmConfig.validate. Mirrors liter-llm’s ChatCompletionRequest.frequency_penalty.
reasoning_effort const char* NULL Reasoning effort level for extended-thinking models, applied to individual requests built from this config. Mirrors liter-llm’s ChatCompletionRequest.reasoning_effort (types.chat.ReasoningEffort). A request-time parameter like temperature/ max_tokens above, not a client-level setting — into_client_builder does not map it. Accepted as a plain string — one of "low", "medium", "high", "minimal", "max" (case-insensitive; liter-llm’s own #[serde(rename_all = "lowercase")] spelling) — rather than importing liter-llm’s enum, because this module compiles even when the liter-llm feature is disabled. See llm.client.parse_reasoning_effort for the conversion into liter_llm.ReasoningEffort.
extra_body const char* NULL Provider-specific extra parameters merged into the request body (guardrails, safety settings, grounding config, etc.), applied to individual requests built from this config. Mirrors liter-llm’s ChatCompletionRequest.extra_body. A request-time parameter like temperature/max_tokens above, not a client-level setting.
load_env int32_t* NULL Whether liter-llm should load provider credentials from environment variables. Mirrors liter-llm’s ClientConfigBuilder.load_env. When NULL, liter-llm’s own default behavior applies.
headers const char* NULL Extra HTTP headers sent with every request to the provider. Mirrors liter-llm’s ClientConfigBuilder.header, for gateways or providers that require custom auth/routing headers.
providers const char* NULL Custom provider configurations, in addition to liter-llm’s built-in providers. Mirrors liter-llm’s LlmConfig.providers, for OpenAI-compatible gateways and self-hosted model servers that are not in the built-in provider catalog.
cache XBERGAlefHandle NULL Response cache configuration. Mirrors liter-llm’s LlmConfig.cache. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value is accepted but unused. Boxed for the same reason as bedrock (a4579589ac): LlmConfig is the payload of EmbeddingModelType.Llm and RerankerModelType.Llm, whose other variants are tens of bytes. Inlining this and the two sub-configs below pushed that variant to 480 bytes and tripped clippy.large_enum_variant on the --features full leg.
budget XBERGAlefHandle NULL Budget enforcement configuration. Mirrors liter-llm’s LlmConfig.budget. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value is accepted but unused.
rate_limit XBERGAlefHandle NULL Per-model rate limiting configuration. Mirrors liter-llm’s LlmConfig.rate_limit. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value is accepted but unused.
cost_tracking int32_t* NULL Enable per-request cost tracking. Mirrors liter-llm’s LlmConfig.cost_tracking. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value is accepted but unused.
tracing int32_t* NULL Enable OpenTelemetry-compatible tracing spans. Mirrors liter-llm’s LlmConfig.tracing. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value is accepted but unused.
cooldown_secs uint64_t* NULL Cooldown duration after transient errors, in seconds. Mirrors liter-llm’s LlmConfig.cooldown_secs. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value is accepted but unused.
health_check_secs uint64_t* NULL Background health check interval, in seconds. Mirrors liter-llm’s LlmConfig.health_check_secs. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value is accepted but unused.
bedrock XBERGAlefHandle NULL AWS Bedrock settings (region, cross-region routing, explicit credentials). Only consulted for bedrock/-prefixed models. When NULL — or when an individual field inside it is NULL — liter-llm falls back to the standard AWS environment variables and the default credential chain.
credential_provider XBERGAlefHandle NULL Managed OAuth2/STS credential provider for auth modes liter-llm cannot express via a static api_key — Azure AD, Vertex AI OAuth2, Vertex AI Application Default Credentials, and AWS STS Web Identity (EKS IRSA) for Bedrock. Mirrors liter-llm’s client.ClientConfigBuilder.credential_provider, which takes an Arc<dyn liter_llm.auth.CredentialProvider> trait object — that cannot appear in a serde DTO. Every CredentialProviderConfig variant is plain data instead, so it round-trips through TOML/JSON/YAML and every language binding like the rest of LlmConfig. Managed credential providers are unavailable on wasm32, where liter-llm uses browser HTTP rather than its native authentication modules. LlmConfig.validate rejects a configured provider on that target instead of silently ignoring it. GitHub Copilot’s device-flow provider has no variant here: it takes no configuration at all (liter_llm.auth.github_copilot.GithubCopilotCredentialProvider.new accepts only an HTTP client) and drives an interactive terminal prompt, so it cannot be expressed as data. A Rust embedder who needs it — or any other fully custom CredentialProvider — can call xberg.llm.client.create_client_with_credential_provider directly with a liter-llm dependency of their own.
max_concurrency uintptr_t* NULL Maximum number of simultaneously in-flight requests to the LLM provider this config resolves to. This is a real, global bound on provider concurrency, not a per-extraction allowance: xberg.llm.client.create_client shares one process-wide client instance per distinct resolved config, so every concurrent extraction that resolves to the same config shares the one in-flight-request limit this value configures, instead of each minting its own (GH#1465). NULL means unlimited. PDF and image OCR batch sizing are not derived from this field, even when the configured OCR backend or vlm_fallback policy can reach a VLM — those call sites mix CPU-bound raster/OCR work with, at most, occasional remote requests, so they size their batches from the general thread budget (max_threads) unconditionally (GH#1465). Captioning is the one feature that additionally uses this value to bound its own per-extraction async request fan-out (issuing only VLM requests, with no CPU batching to protect), on top of the global provider-side limit described above; that per-extraction bound clamps a value below 1 up to 1. The global provider-side limit does not clamp: Some(0) reaches liter-llm as-is, and liter-llm rejects it when building the client (zero permitted in-flight requests is never useful), so it surfaces as a create_client error rather than a silent clamp to 1. This field is intentionally last to preserve positional constructor compatibility in generated language bindings.
max_response_bytes uintptr_t* NULL Maximum size, in bytes, of a single HTTP response body read from the LLM provider. Bounds response bodies on every non-streaming call (chat completions, embeddings, model listings, …) and the error body read on a failed request; a successful streaming response keeps its own existing frame bounds and is unaffected. NULL (the default) means unbounded, matching liter-llm’s own default. Mirrors liter-llm’s client.ClientConfigBuilder.max_response_bytes, which is native-only (native-http, non-wasm32) — see llm.client.build_client_config. Some(0) is rejected by LlmConfig.validate rather than reaching liter-llm, which would refuse it at client-build time with the same complaint.

Since: v1.1

C representation: XBERGLlmProviderConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLlmProviderConfig does not appear anywhere in the generated header.

A custom provider configuration entry, in addition to liter-llm’s built-in providers.

Mirrors liter-llm’s LlmProviderConfig.

Field Type Default Description
name const char* — Provider name, used to key model prefix matching.
base_url const char* — Base URL for the provider’s OpenAI-compatible API.
auth_header const char* NULL Header name used to carry the API key (defaults to Authorization when unset).
model_prefixes const char* NULL Model name prefixes routed to this provider (e.g. ["my-provider/"]).

Since: v1.1

C representation: XBERGLlmRateLimitConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLlmRateLimitConfig does not appear anywhere in the generated header.

Per-model rate limiting configuration.

Mirrors liter-llm’s LlmRateLimitConfig. Only takes effect when liter-llm’s tower feature is compiled in; otherwise the value round-trips through configuration but is not enforced at request time.

Field Type Default Description
rpm uint32_t* NULL Requests per minute limit.
tpm uint64_t* NULL Tokens per minute limit.
window_seconds uint64_t* NULL Rate limit window, in seconds.

C representation: XBERGLlmUsage is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGLlmUsage does not appear anywhere in the generated header.

Token usage and cost data for a single LLM call made during extraction.

Populated when VLM OCR, structured extraction, or LLM-based embeddings are used. Multiple entries may be present when multiple LLM calls occur within one extraction (e.g. VLM OCR + structured extraction).

Field Type Default Description
model const char* — The LLM model identifier (e.g. “openai/gpt-4o”, “anthropic/claude-sonnet-4-20250514”).
source const char* — The pipeline stage that triggered this LLM call (e.g. “vlm_ocr”, “structured_extraction”, “embeddings”).
input_tokens uint64_t* NULL Number of input/prompt tokens consumed.
output_tokens uint64_t* NULL Number of output/completion tokens generated.
total_tokens uint64_t* NULL Total tokens (input + output).
estimated_cost double* NULL Estimated cost in USD based on the provider’s published pricing.
finish_reason const char* NULL Why the model stopped generating (e.g. “stop”, “length”, “content_filter”).

C representation: XBERGMapResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGMapResult does not appear anywhere in the generated header.

The result of a map operation, containing discovered URLs.

Field Type Default Description
urls const char* NULL The list of discovered URLs.

C representation: XBERGMarkdownCodeBlock is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGMarkdownCodeBlock does not appear anywhere in the generated header.

A fenced code block extracted from Markdown.

Field Type Default Description
language const char* — Declared language identifier, or an empty string when absent.
code const char* — Code block content.

C representation: XBERGMarkdownLink is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGMarkdownLink does not appear anywhere in the generated header.

A link extracted from Markdown.

Field Type Default Description
text const char* — Visible link text.
url const char* — Link destination.

C representation: XBERGMetaSchema is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGMetaSchema does not appear anywhere in the generated header.

Compiled meta-schema validator over preset.schema.json.

Compile the given JSON text as a Draft 2020-12 meta-schema.

Signature:

XBERGAlefHandle xberg_meta_schema_compile(const char* meta_schema_json);

Example:

XBERGAlefHandle result = xberg_meta_schema_compile("value");

Parameters:

Name Type Required Description
meta_schema_json const char* Yes The meta schema json

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.

Validate raw against the meta-schema and deserialize into a Preset, stamping the fingerprint over the canonical file bytes.

Signature:

XBERGAlefHandle xberg_meta_schema_parse_preset(XBERGAlefHandle this, const char* path, const uint8_t* raw);

Example:

XBERGAlefHandle result = xberg_meta_schema_parse_preset(instance, "value", (const uint8_t *)"data");

Parameters:

Name Type Required Description
path const char* Yes Path to the file
raw const uint8_t* Yes The raw

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.


C representation: XBERGMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGMetadata does not appear anywhere in the generated header.

Extraction result metadata.

Contains common fields applicable to all formats, format-specific metadata via a discriminated union, and additional custom fields from postprocessors.

Field Type Default Description
title const char* NULL Document title
subject const char* NULL Document subject or description
authors const char* NULL Primary author(s) - always Vec for consistency
keywords const char* NULL Keywords/tags - always Vec for consistency
language const char* NULL Primary language (ISO 639 code)
created_at const char* NULL Creation timestamp (ISO 8601 format)
modified_at const char* NULL Last modification timestamp (ISO 8601 format)
created_by const char* NULL User who created the document
modified_by const char* NULL User who last modified the document
pages XBERGAlefHandle NULL Page/slide/sheet structure with boundaries
format XBERGAlefHandle NULL Format-specific metadata (discriminated union) Contains detailed metadata specific to the document format. Serialized as a nested "format" object with a format_type discriminator field.
image_preprocessing XBERGAlefHandle NULL Image preprocessing metadata (when OCR preprocessing was applied)
json_schema const char* NULL JSON schema (for structured data extraction)
error XBERGAlefHandle NULL Error metadata (for batch operations)
extraction_duration_ms uint64_t* NULL Extraction duration in milliseconds (for benchmarking). This field is populated by batch extraction to provide per-file timing information. It’s NULL for single-file extraction (which uses external timing).
category const char* NULL Document category (from frontmatter or classification).
tags const char* NULL Document tags (from frontmatter).
document_version const char* NULL Document version string (from frontmatter).
abstract_text const char* NULL Abstract or summary text (from frontmatter).
output_format const char* NULL Output format identifier (e.g., “markdown”, “html”, “text”). Set by the output format pipeline stage when format conversion is applied. Previously stored in metadata.additional["output_format"].
ocr_used int32_t — Whether OCR was used during extraction. Set to true whenever the extraction pipeline ran an OCR backend (Tesseract, PaddleOCR, VLM, etc.) and used that output as the primary or fallback text. false means native text extraction was used exclusively.
additional const char* NULL Additional custom fields from postprocessors. Serialized as a nested "additional" object (not flattened at root level). Uses Cow<'static, str> keys so static string keys avoid allocation.

C representation: XBERGModelPaths is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGModelPaths does not appear anywhere in the generated header.

Combined paths to all models needed for OCR (backward compatibility).

Field Type Default Description
det_model const char* — Exact path to the detection ONNX model in the Hugging Face snapshot.
cls_model const char* — Exact path to the classification ONNX model in the Hugging Face snapshot.
rec_model const char* — Exact path to the recognition ONNX model in the Hugging Face snapshot.
dict_file const char* — Path to the character dictionary file.

C representation: XBERGMultiVectorEmbedding is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGMultiVectorEmbedding does not appear anywhere in the generated header.

A ColBERT multi-vector embedding: one row per attention-live token.

data is a flat, row-major buffer of length num_tokens * dim — row i (the embedding for token i) occupies data[i*dim .. (i+1)*dim]. Flat storage keeps the type FFI-friendly across binding boundaries; use MultiVectorEmbedding.rows internally to iterate per-token slices.

Field Type Default Description
num_tokens uint32_t — Number of attention-live token rows (padding rows are dropped, not zeroed — see engine.normalize_tokens).
dim uint32_t — Dimensionality of each per-token vector.
data const char* — Flat row-major buffer, length num_tokens * dim.

C representation: XBERGMultidocInput is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGMultidocInput does not appear anywhere in the generated header.

Input signals for multi-document boundary detection.

Field Type Default Description
page_count uint32_t — Total number of pages in the PDF.
pages const char* — Per-page signals extracted from the PDF.

C representation: XBERGMultidocThresholds is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGMultidocThresholds does not appear anywhere in the generated header.

Thresholds for multi-document boundary detection.

All fields are public; callers override any subset via struct-update syntax.

Field Type Default Description
density_shift_threshold float 0.3 Text density difference threshold for DensityShift detection. Default: 0.3.
bigram_overlap_min float 0.1 Minimum bigram-overlap ratio below which a density shift is promoted to a DensityShift boundary. Default: 0.1 (10 % overlap).

Since: v1.0

C representation: XBERGNerConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGNerConfig does not appear anywhere in the generated header.

Configuration for the NER post-processor.

Field Type Default Description
backend XBERGAlefHandle XBERG_ONNX Backend that runs the entity detection.
categories const char* NULL Entity categories to detect. Defaults to a sensible PERSON/ORG/LOCATION/EMAIL set when empty.
model const char* NULL Override the default model — only used by NerBackendKind.Onnx. NULL lets the backend pick its pinned default xberg GLiNER model alias.
llm XBERGAlefHandle NULL Optional LLM configuration — only used by NerBackendKind.Llm. Token usage for LLM backends is recorded in ExtractedDocument.llm_usage.
custom_labels const char* NULL Arbitrary user-supplied entity labels for zero-shot detection. xberg-gliner natively supports zero-shot inference over caller-supplied labels. The LLM backend also honours these labels by including them in the structured-output schema. Custom labels surface as EntityCategory.Custom in the resulting Entity stream. Use this when you need domain-specific entity types (e.g. "Treatment", "Product", "Vessel") without forking GLiNER’s taxonomy.

C representation: XBERGNgramRange is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGNgramRange does not appear anywhere in the generated header.

Inclusive word-count range used to form keyword candidates.

Field Type Default Description
min uintptr_t 1 Minimum number of words in a candidate.
max uintptr_t 3 Maximum number of words in a candidate.

C representation: XBERGOcrBackend is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrBackend does not appear anywhere in the generated header.

Trait for OCR backend plugins.

Implement this trait to add custom OCR capabilities. OCR backends can be:

  • Native Rust implementations (like Tesseract)
  • FFI bridges to external libraries (like PaddleOCR)
  • Cloud-based OCR services (Google Vision, AWS Textract, etc.)

OCR backends must be thread-safe (Send + Sync) to support concurrent processing.

Process an image and extract text via OCR.

Returns:

An ExtractedDocument containing the extracted text and metadata.

Errors:

  • XbergError.Ocr - OCR processing failed
  • XbergError.Validation - Invalid image format or configuration
  • XbergError.Io - I/O errors (these always bubble up)

Backends that support runtime tuning can read config.backend_options and deserialize only the keys they care about. Unknown keys are silently ignored, so multiple backends can coexist in a pipeline without key conflicts.

Signature:

XBERGAlefHandle xberg_ocr_backend_process_image(XBERGAlefHandle this, const uint8_t* image_bytes, XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_ocr_backend_process_image(instance, (const uint8_t *)"data", 0);

Parameters:

Name Type Required Description
image_bytes const uint8_t* Yes Raw image data (JPEG, PNG, TIFF, etc.)
config XBERGAlefHandle Yes OCR configuration (language, PSM mode, etc.)

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.

Process a file and extract text via OCR.

Default implementation reads the file and calls process_image. Override for custom file handling or optimizations.

Errors:

Same as process_image, plus file I/O errors.

Signature:

XBERGAlefHandle xberg_ocr_backend_process_image_file(XBERGAlefHandle this, const char* path, XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_ocr_backend_process_image_file(instance, "value", 0);

Parameters:

Name Type Required Description
path const char* Yes Path to the image file
config XBERGAlefHandle Yes OCR configuration

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.

Check if this backend supports a given language code.

Returns:

true if the language is supported, false otherwise.

Signature:

int32_t xberg_ocr_backend_supports_language(XBERGAlefHandle this, const char* lang);

Example:

int32_t result = xberg_ocr_backend_supports_language(instance, "value");

Parameters:

Name Type Required Description
lang const char* Yes ISO 639-2/3 language code (e.g., “eng”, “deu”, “fra”)

Returns: int32_t

Get the backend type identifier.

Returns:

The backend type enum value.

Signature:

XBERGAlefHandle xberg_ocr_backend_backend_type(XBERGAlefHandle this);

Example:

XBERGAlefHandle result = xberg_ocr_backend_backend_type(instance);

Returns: XBERGAlefHandle

Optional: Get a list of all supported languages.

Defaults to empty list. Override to provide comprehensive language support info.

Signature:

const char* xberg_ocr_backend_supported_languages(XBERGAlefHandle this);

Example:

const char* result = xberg_ocr_backend_supported_languages(instance);

Returns: const char*

xberg_ocr_backend_supported_languages_for()
Section titled “xberg_ocr_backend_supported_languages_for()”

Optional: languages usable under config, defaulting to ignoring it. GH#1857.

Signature:

const char* xberg_ocr_backend_supported_languages_for(XBERGAlefHandle this, XBERGAlefHandle config);

Example:

const char* result = xberg_ocr_backend_supported_languages_for(instance, 0);

Parameters:

Name Type Required Description
config XBERGAlefHandle Yes The configuration options

Returns: const char*

Optional: whether language is supported under config. GH#1857.

Signature:

int32_t xberg_ocr_backend_supports_language_for(XBERGAlefHandle this, XBERGAlefHandle config, const char* language);

Example:

int32_t result = xberg_ocr_backend_supports_language_for(instance, 0, "value");

Parameters:

Name Type Required Description
config XBERGAlefHandle Yes The configuration options
language const char* Yes The language

Returns: int32_t

xberg_ocr_backend_supports_table_detection()
Section titled “xberg_ocr_backend_supports_table_detection()”

Optional: Check if the backend supports table detection.

Defaults to false. Override if your backend can detect and extract tables.

Signature:

int32_t xberg_ocr_backend_supports_table_detection(XBERGAlefHandle this);

Example:

int32_t result = xberg_ocr_backend_supports_table_detection(instance);

Returns: int32_t

xberg_ocr_backend_supports_document_processing()
Section titled “xberg_ocr_backend_supports_document_processing()”

Check if the backend supports direct document-level processing (e.g. for PDFs).

Defaults to false. Override if the backend has optimized document processing. PDF extraction uses this optimized path only when both effective page margins are zero; nonzero margins require per-page image processing so geometry can be filtered correctly.

Signature:

int32_t xberg_ocr_backend_supports_document_processing(XBERGAlefHandle this);

Example:

int32_t result = xberg_ocr_backend_supports_document_processing(instance);

Returns: int32_t

xberg_ocr_backend_emits_structured_markdown()
Section titled “xberg_ocr_backend_emits_structured_markdown()”

Declare that this backend emits structured markdown directly (tables, headings, lists) and downstream layout reconstruction should be skipped.

Defaults to false — classical OCR backends (Tesseract, PaddleOCR classical) return plain text per detected region. End-to-end VLM backends (PaddleOCR-VL, GOT-OCR 2.0) emit markdown in one forward pass and should override this to true.

Signature:

int32_t xberg_ocr_backend_emits_structured_markdown(XBERGAlefHandle this);

Example:

int32_t result = xberg_ocr_backend_emits_structured_markdown(instance);

Returns: int32_t

Declare how this backend’s reported page-level confidence must be interpreted.

Defaults to ConfidenceSemantics.Uncalibrated. This default is deliberately the least trusting option, not ConfidenceSemantics.Legibility: a new backend that reports some confidence number is not thereby safe to gate on, and defaulting to Legibility would let the next backend silently inherit a threshold calibrated for a different backend’s scale — exactly the failure this type exists to prevent (see the type’s doc comment). Override this only after validating that the reported number tracks legibility on a known scale.

Signature:

XBERGAlefHandle xberg_ocr_backend_confidence_semantics(XBERGAlefHandle this);

Example:

XBERGAlefHandle result = xberg_ocr_backend_confidence_semantics(instance);

Returns: XBERGAlefHandle

xberg_ocr_backend_page_orientation_handling()
Section titled “xberg_ocr_backend_page_orientation_handling()”

Declare how this backend copes with a page raster whose text is not upright.

Defaults to PageOrientationHandling.RequiresUpright. This default is deliberately the least capable option, not PageOrientationHandling.SelfCorrecting: a new backend must not silently inherit Tesseract’s ability to reconstruct reading order on a rotated raster and then quietly emit garbage the first time it faces one (see the type’s doc comment for the measured A/B behind this). Override this only after validating the backend’s actual behaviour on a rotated page.

Signature:

XBERGAlefHandle xberg_ocr_backend_page_orientation_handling(XBERGAlefHandle this);

Example:

XBERGAlefHandle result = xberg_ocr_backend_page_orientation_handling(instance);

Returns: XBERGAlefHandle

Process a document file directly via OCR.

Only called if supports_document_processing returns true.

Signature:

XBERGAlefHandle xberg_ocr_backend_process_document(XBERGAlefHandle this, const char* path, XBERGAlefHandle config);

Example:

XBERGAlefHandle result = xberg_ocr_backend_process_document(instance, "value", 0);

Parameters:

Name Type Required Description
path const char* Yes The path
config XBERGAlefHandle Yes The ocr config

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.


C representation: XBERGOcrBackendCapabilities is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrBackendCapabilities does not appear anywhere in the generated header.

A registered OCR backend’s declared name and language capabilities.

Returned by list_ocr_backend_capabilities. See that function’s documentation for the determinism guarantees and the important caveat about what an empty supported_languages means.

Field Type Default Description
name const char* — The backend’s registered name, as returned by list_ocr_backends.
supported_languages const char* — The languages this backend declares support for, via OcrBackend.supported_languages. An empty list means the backend does not enumerate its languages — it is not a statement that the backend supports no languages. OcrBackend.supported_languages is a defaulted trait method that returns [], and not every backend overrides it: the VLM backend (llm.vlm_ocr.VlmOcrBackend) accepts every language via supports_language while inheriting the empty default here. Use ocr_backend_supports_language to decide whether one specific language is usable — never infer “unsupported” from an empty list. The order of this list is preserved exactly as the backend reported it and is not re-sorted. Tesseract’s order comes from enumerating installed tessdata files; PaddleOCR’s comes from its own SUPPORTED_LANGUAGES constant. Re-sorting would disagree with the precedence each backend’s own supports_language implementation uses internally. This is the no-override answer (list_ocr_backend_capabilities). Use list_ocr_backend_capabilities_for when the caller sets OcrConfig.tessdata_path: for Tesseract the language list is a property of the resolved tessdata directory, and the no-override chain can name a different directory than the one a job using that config will load from. See GH#1857.

C representation: XBERGOcrConfidence is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrConfidence does not appear anywhere in the generated header.

Confidence scores for an OCR element.

Separates detection confidence (how confident that text exists at this location) from recognition confidence (how confident about the actual text content).

Field Type Default Description
detection double* NULL Detection confidence: how confident the OCR engine is that text exists here. PaddleOCR provides this as box_score, Tesseract doesn’t have a direct equivalent. Range: 0.0 to 1.0 (or None if not available).
recognition double — Recognition confidence: how confident about the text content. Range: 0.0 to 1.0.

C representation: XBERGOcrConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrConfig does not appear anywhere in the generated header.

OCR configuration.

Field Type Default Description
enabled int32_t true Whether OCR is enabled. Setting enabled: false is a shorthand for disable_ocr: true on the parent ExtractionConfig. Images return metadata only; PDFs use native text extraction without OCR fallback. Defaults to true. When false, all other OCR settings are ignored.
backend const char* "tesseract" OCR backend: tesseract, paddleocr, paddle-ocr, sceptre, or vlm. Sceptre uses ONNX Runtime on desktop/server and tract on supported mobile builds. Browser WebAssembly uses the separate byte-fed Sceptre worker API.
language const char* ["eng"] Language code(s) for OCR recognition. Defaults to ["eng"]. For Tesseract, languages are joined with “+”. A list is the canonical form and the only form accepted by the binding object APIs (Python, Node, PHP, WASM, etc.): ["eng", "deu"]. When deserializing from a config file, JSON body, or the REST/MCP API, a single string is also accepted, either as one code (“eng”) or “+”-joined (“eng+deu”). The four candle-based backends also use this list to decide which scripts are plausible in their output, dropping a line written in an unconfigured script.
tesseract_config XBERGAlefHandle NULL Tesseract-specific configuration (optional)
output_format XBERGAlefHandle NULL Output format for OCR results (optional, for format conversion)
paddle_ocr_config const char* NULL PaddleOCR-specific configuration (optional, JSON passthrough). Deserialized into a PaddleOcrConfig, so any of its fields can be overridden here — most notably model_version ("pp-ocrv6" default / "pp-ocrv5") and model_tier. In TOML: toml [ocr.paddle_ocr_config] model_version = "pp-ocrv5" model_tier = "server" The XBERG_OCR_MODEL_VERSION / XBERG_OCR_MODEL_TIER environment variables set the same two keys for env-configured servers (issue #1279).
backend_options const char* NULL Arbitrary per-call options passed through to the backend unchanged. Custom OCR backends and built-in backends that support runtime tuning can read this value and deserialize the keys they care about. Keys unknown to the backend are silently ignored. This is the recommended extension point for per-call parameters that are not covered by the typed fields above (e.g. mode switching, preprocessing flags, inference batch size). Scope: when pipeline is NULL, this value is propagated to the primary stage of the auto-constructed pipeline. When pipeline is explicitly set, this field has no effect — the caller must set OcrPipelineStage.backend_options directly on the relevant stage(s) instead. Example: json { "mode": "fast", "enable_layout": true, "timeout_ms": 5000 }
element_config XBERGAlefHandle NULL OCR element extraction configuration
quality_thresholds XBERGAlefHandle NULL Quality thresholds for the native-text-to-OCR fallback decision. When None, uses compiled defaults (matching previous hardcoded behavior).
pipeline XBERGAlefHandle NULL Multi-backend OCR pipeline configuration. When set, enables weighted fallback across multiple OCR backends based on output quality. When None, uses the single backend field (same as today).
auto_rotate int32_t false Enable automatic page rotation based on orientation detection. When enabled, page orientation (0/90/180/270 degrees) is detected with an ONNX PP-LCNet document-orientation classifier — NOT Tesseract’s own DetectOrientationScript()/OSD, which this library does not call. If the page is rotated with high confidence, the image is corrected before recognition. Applies to every OCR backend; the tesseract-only TesseractConfig.preprocessing field of the same name is OR’d in on top of this one and affects tesseract alone. Independent of a page’s PDF /Rotate entry, which is handled separately on every OCR route regardless of this setting.
vlm_fallback XBERGAlefHandle XBERG_DISABLED Ergonomic VLM fallback policy. When set to anything other than VlmFallbackPolicy.Disabled and OcrConfig.pipeline is NULL, a multi-stage pipeline is synthesised automatically: - VlmFallbackPolicy.OnLowQuality → [classical_stage, vlm_stage] with the quality_threshold mapped onto OcrQualityThresholds.pipeline_min_quality. - VlmFallbackPolicy.Always → [vlm_stage] only. Requires OcrConfig.vlm_config to be Some when not Disabled. When OcrConfig.pipeline is explicitly set, this field is ignored.
vlm_config XBERGAlefHandle NULL VLM (Vision Language Model) OCR configuration. Required when backend is "vlm" or when vlm_fallback is not VlmFallbackPolicy.Disabled. Uses liter-llm to send page images to a vision model for text extraction.
vlm_prompt const char* NULL Custom Jinja2 prompt template for VLM OCR. When NULL, uses the default template. Available variables: - {{ language }} — The document language code (e.g., “eng”, “deu”).
acceleration XBERGAlefHandle NULL Hardware acceleration for ONNX Runtime models (e.g. PaddleOCR, layout detection). Not user-configurable via config files — injected at runtime from ExtractionConfig.acceleration before each process_image call.
security_limits XBERGAlefHandle NULL Security limits applied when decoding raw image bytes for OCR (GH#1554). Not user-configurable via config files — injected at runtime from ExtractionConfig.security_limits before each process_image call, the same pattern Self.acceleration uses. ExtractionConfig.security_limits remains the source of truth and overwrites this field whenever it carries a value; a value set here directly is honoured only when ExtractionConfig carries none, so the two cannot drift while a directly-set limit is no longer silently discarded (GH#1651). A backend consulting a backend_options override for this call may still let that override win, but in the absence of one this field is what backends should fall back to instead of SecurityLimits.default(). NULL means “use SecurityLimits.default()”, never “disable the check”.
tessdata_bytes const char* NULL Caller-supplied Tesseract traineddata bytes per language code. Primary use case is the WASM build, which has no filesystem and cannot download tessdata at runtime. Native builds typically rely on TessdataManager and ignore this field. When present, the WASM Tesseract backend prefers these bytes over its compile-time-bundled English data. Skipped by serde to keep config files small — supply via the typed API at runtime.
tessdata_path const char* NULL Runtime override for tessdata directory path. When set, uses this path as the highest-priority tessdata location, bypassing environment variables and cache directories. Useful for embedding pre-installed tessdata in applications. When NULL, uses the standard resolution chain: TESSDATA_PREFIX env, cache dir, system paths.
numeric_repair int32_t false Repair OCR tokens that are clearly numeric but mis-punctuated: a dropped thousands separator, a decimal point misread for a grouping comma, or one number split into two tokens at a rendering gap (GH#1789). Defaults to false. Unlike the always-on list-marker repair (crate.extractors.pdf.ocr.scoring.repair_ocr_list_markers), this repair has no table-column context available at the point OCR text comes back as a flat string, so it cannot tell "1.234,56" (European) from "1,234.56" (US) apart on its own – it always assumes the US/UK convention (comma groups, period decimals). The rules read only the punctuation and digit counts, never whether a token is an amount, so each of these is rewritten too (GH#1836). Enable it only for documents whose numbers are amounts, and which use that convention: - A bare 4-9 digit integer becomes grouped, whatever it denotes: "Year 2019" -> "Year 2,019", and likewise a ZIP code, a page number or a part number. Only a literal "FY " prefix is exempt. - A US-convention decimal with exactly three fraction digits and at most three integer digits becomes an amount: "0.125" -> "0,125". Four fraction digits ("0.7906") or a trailing % are left alone. - A lone digit followed by a single space and an already-grouped number is joined, which is indistinguishable from two adjacent table cells: "5 12,000" -> "512,000". The repair is skipped entirely when tesseract_config.output_format is "hocr" or "tsv", because it would re-punctuate the bare coordinate integers in that markup.

C representation: XBERGOcrElement is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrElement does not appear anywhere in the generated header.

A unified OCR element representing detected text with full metadata.

This is the primary type for structured OCR output, preserving all information from both Tesseract and PaddleOCR backends.

Field Type Default Description
text const char* — The recognized text content.
geometry XBERGAlefHandle XBERG_RECTANGLE Bounding geometry (rectangle or quadrilateral).
confidence XBERGAlefHandle — Confidence scores for detection and recognition.
level XBERGAlefHandle XBERG_LINE Hierarchical level (word, line, block, page).
rotation XBERGAlefHandle NULL Rotation information (if detected).
page_number uint32_t 1 Page number (1-indexed).
parent_id const char* NULL Parent element ID for hierarchical relationships. When hierarchy output is enabled, this resolves to another emitted element’s backend_metadata["element_id"] value.
backend_metadata const char* NULL Backend-specific metadata that doesn’t fit the unified schema.

C representation: XBERGOcrElementConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrElementConfig does not appear anywhere in the generated header.

Configuration for OCR element extraction.

Controls how OCR elements are extracted and filtered.

Field Type Default Description
include_elements int32_t — Whether to include OCR elements in the extraction result. When true, the ocr_elements field in ExtractedDocument will be populated.
min_level XBERGAlefHandle XBERG_LINE Minimum hierarchical level to include. Elements below this level (e.g., words when min_level is Line) will be excluded.
min_confidence double — Minimum recognition confidence threshold (0.0-1.0). Elements with confidence below this threshold will be filtered out.
build_hierarchy int32_t — Whether to build hierarchical relationships between elements. When true, emitted elements receive an element_id metadata value and parent_id references are populated only when a spatially containing parent is also emitted.

C representation: XBERGOcrExtractionResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrExtractionResult does not appear anywhere in the generated header.

OCR extraction result.

Result of performing OCR on an image or scanned document, including recognized text and detected tables.

Field Type Default Description
content const char* — Recognized text content
mime_type const char* — Original MIME type of the processed image
metadata const char* NULL OCR processing metadata (confidence scores, language, etc.)
tables const char* NULL Tables detected and extracted via OCR
ocr_elements const char* NULL Structured OCR elements with bounding boxes and confidence scores. Available when TSV output is requested or table detection is enabled.

C representation: XBERGOcrMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrMetadata does not appear anywhere in the generated header.

OCR processing metadata.

Captures information about OCR processing configuration and results.

Field Type Default Description
language const char* — OCR language code(s) used
psm int32_t — Tesseract Page Segmentation Mode (PSM)
output_format const char* — Output format (e.g., “text”, “hocr”)
table_count uint32_t — Number of tables detected
table_rows uint32_t* NULL Number of rows in the detected table (if a single table was found).
table_cols uint32_t* NULL Number of columns in the detected table (if a single table was found).

C representation: XBERGOcrPipelineConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrPipelineConfig does not appear anywhere in the generated header.

Multi-backend OCR pipeline with quality-based fallback.

Backends are tried in priority order (highest first). After each backend produces output, quality is evaluated. If it meets quality_thresholds.pipeline_min_quality, the result is accepted. Otherwise the next backend is tried; if none clears the threshold, an internal selection policy derived from the OcrConfig decides which stage’s result is returned as the best effort (vlm_fallback pipelines prefer their last non-empty stage; explicit and classical pipelines stay score-based).

Field Type Default Description
stages const char* — Ordered list of backends to try. Sorted by priority (descending) at runtime.
quality_thresholds XBERGAlefHandle /* serde(default) */ Quality thresholds for deciding whether to accept a result or try the next backend.

C representation: XBERGOcrPipelineStage is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrPipelineStage does not appear anywhere in the generated header.

A single backend stage in the OCR pipeline.

Field Type Default Description
backend const char* — Backend name: “tesseract”, “paddleocr”, “paddle-ocr”, “sceptre”, “vlm”, or a custom registered name. Sceptre uses ONNX Runtime on desktop/server and tract on Android/iOS; browser WebAssembly has a separate byte-fed engine because the normal async OCR registry assumes native model storage.
priority uint32_t 100 Priority weight (higher = tried first). Stages are sorted by priority descending.
language const char* /* serde(default) */ Language override for this stage (None = use parent OcrConfig.language). A list is the canonical form and the only form accepted by the binding object APIs: ["eng", "deu"]. When deserializing from a config file, JSON body, or the REST/MCP API, a single string is also accepted, either as one code (“eng”) or “+”-joined (“eng+deu”).
tesseract_config XBERGAlefHandle /* serde(default) */ Tesseract-specific config override for this stage.
paddle_ocr_config const char* /* serde(default) */ PaddleOCR-specific config for this stage.
vlm_config XBERGAlefHandle /* serde(default) */ VLM config override for this pipeline stage.
backend_options const char* /* serde(default) */ Arbitrary per-call options passed through to the backend unchanged. Backends that support runtime tuning (mode switching, preprocessing flags, inference parameters, etc.) read this value and deserialize the keys they care about. Keys unknown to the backend are silently ignored, so options from different backends can coexist in the same config without conflict. Example (custom backend): json { "mode": "fast", "enable_layout": true }

C representation: XBERGOcrPoint is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrPoint does not appear anywhere in the generated header.

A point in OCR raster pixel coordinates.

Field Type Default Description
x uint32_t — Horizontal coordinate.
y uint32_t — Vertical coordinate.

C representation: XBERGOcrQualityThresholds is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrQualityThresholds does not appear anywhere in the generated header.

Quality thresholds for OCR fallback decisions and pipeline quality gating.

Fields default to conservative extraction behavior. Suspected OCR recognition noise is reported but retained unless destructive filtering is explicitly enabled.

Field Type Default Description
min_total_non_whitespace uintptr_t 64 Minimum total non-whitespace characters to consider text substantive.
min_non_whitespace_per_page double 32 Minimum non-whitespace characters per page on average.
min_meaningful_word_len uintptr_t 4 Minimum character count for a word to be “meaningful”.
min_meaningful_words uintptr_t 3 Minimum count of meaningful words before text is accepted.
min_alnum_ratio double 0.3 Minimum alphanumeric ratio (non-whitespace chars that are alphanumeric).
min_garbage_chars uintptr_t 5 Minimum Unicode replacement characters (U+FFFD) to trigger OCR fallback.
max_fragmented_word_ratio double 0.6 Maximum fraction of short (1-2 char) words before text is considered fragmented.
critical_fragmented_word_ratio double 0.8 Critical fragmentation threshold — triggers OCR regardless of meaningful words. Normal English text has ~20-30% short words. 80%+ is definitive garbage.
max_ocr_output_fragmented_word_ratio double 0.35 Maximum fraction of short (1-2 char) words an OCR result may carry before the page is reported as suspected recognition noise. This is a different decision, and a different operating point, from Self.max_fragmented_word_ratio / Self.critical_fragmented_word_ratio: those ask “is the native text bad enough that we should OCR this page?”, and are tuned to only fire on text that is already definitively broken (0.6 / 0.8). Here we are judging what OCR produced, where the failure mode is an engine run over line art — a scanned plat, an engineering drawing, a signature flourish — returning confident-looking strings that are not words. Measured over the 16 pages of a recorded municipal ordinance (13 prose pages, 3 scanned survey drawings): prose ran 0.04-0.28, the drawings 0.42-0.47. The default sits in that gap with margin on both sides. By default, crossing the threshold emits a processing warning but retains the recognized text. Set Self.discard_suspected_ocr_noise to true to restore destructive filtering.
min_ocr_mean_confidence double 75 Minimum mean OCR confidence (0-100) below which a page is reported as suspected noise. This is the engine’s own uncertainty about what it read, and it is a far sharper instrument than any statistic derived from the output text. Measured per page over a recorded municipal ordinance (10 prose pages, 6 scanned survey/architectural drawings) with Tesseract 5.5.3: text prose 86.3 - 95.3 drawings 18.5 - 64.3 The default sits in that gap with ~11 points of margin on each side. Compare the short-word ratio, which separated the same two groups by 0.09 on a 0-1 scale. A backend that reports no confidence (no mean_text_conf in its result metadata) skips this check entirely and falls back to Self.max_ocr_output_fragmented_word_ratio. By default, the signal emits a warning without discarding content. Set to 0.0 to disable the signal.
min_words_for_ocr_output_check uintptr_t 20 Minimum word count before Self.max_ocr_output_fragmented_word_ratio may report a page. Short pages (a signature block, an exhibit title) are legitimately dominated by short words, so the ratio is not meaningful on them and the signal stays disabled.
max_ocr_output_dict_invalid_word_ratio double 1.01 Maximum fraction of a Tesseract page’s dictionary-checkable words that TesseractAPI.is_valid_word may reject before the page is treated as recognition noise, supplementing (never replacing) Self.max_ocr_output_fragmented_word_ratio. Tesseract-only: other backends never report this ratio (see dictionary_invalid_word_ratio in ocr.processor.execution), so this threshold is simply never consulted for them. Calibration owed, same as ocr.types.TesseractConfig.min_confidence: unlike Self.max_ocr_output_fragmented_word_ratio (measured across a recorded ordinance’s prose vs. drawing pages), this ratio has not yet had a page-level measurement run over a labeled corpus. The default therefore disables the check entirely (1.01, above the [0.0, 1.0] range a ratio can reach) rather than guess an operating point. Do not lower this without running that measurement first.
discard_suspected_ocr_noise int32_t false Discard non-empty OCR text when any configured recognition-noise signal fires. Defaults to false: suspected noise is retained and surfaced through processing_warnings. Set to true to preserve the legacy destructive behavior.
min_avg_word_length double 2 Minimum average word length. Below this with enough words indicates garbled extraction.
min_words_for_avg_length_check uintptr_t 50 Minimum word count before average word length check applies.
min_consecutive_repeat_ratio double 0.08 Minimum consecutive word repetition ratio to detect column scrambling.
min_words_for_repeat_check uintptr_t 50 Minimum word count before consecutive repetition check is applied.
substantive_min_chars uintptr_t 100 Minimum character count for “substantive markdown” OCR skip gate.
non_text_min_chars uintptr_t 20 Minimum character count for “non-text content” OCR skip gate.
alnum_ws_ratio_threshold double 0.4 Alphanumeric+whitespace ratio threshold for skip decisions.
pipeline_min_quality double 0.5 Minimum quality score (0.0-1.0) for a pipeline stage result to be accepted. If the result from a backend scores below this, try the next backend.
min_undecodable_ratio double 0.5 Minimum fraction of non-whitespace characters that are undecodable (Unicode Private Use Area, replacement characters, or non-whitespace control characters) before a page’s text layer is treated as unreadable and routed to OCR (issue #1254). Gated by min_total_non_whitespace so short snippets with a stray symbol or two do not trip this check.
enable_provenance_ocr_routing int32_t true Whether to route a page to OCR when xberg_native_pdf reports that a high fraction of its text was fabricated rather than read from the file (MappingProvenance.Fallback, xberg_native_pdf 0.3.75+, issue #1254). This is a direct fact from the extractor’s ISO 32000-1 §9.10.2 mapping cascade, distinct from the character-heuristic proxy behind min_undecodable_ratio. Defaults to true.
min_provenance_fallback_ratio double 0.5 Minimum fraction of a page’s non-whitespace characters with MappingProvenance.Fallback provenance before the page is treated as having a fabricated text layer and routed to OCR (issue #1254). Gated by min_total_non_whitespace so a short page with a stray fallback character cannot trip it. Only used when enable_provenance_ocr_routing is true.
enable_plausibility_ocr_routing int32_t true Whether to route a page to OCR when its decoded text does not read as any real, detectable language (issue #1696). Unlike enable_provenance_ocr_routing, this catches a /ToUnicode CMap that resolves every glyph to a character, but consistently the WRONG one (e.g. a ROT-shifted mapping) – text that is structurally indistinguishable from real prose to every character-shape check, including the provenance signal, because the mapping tier really is file-backed. Defaults to true.
min_reliable_language_chunk_ratio double 0.1 Minimum fraction of a page’s language-detection chunks that whatlang classifies as reliable (Info.is_reliable()) before the page is trusted as legible (issue #1696). Below this AND below the mean-confidence guard together, the page’s text layer is treated as implausible and routed to OCR. Only used when enable_plausibility_ocr_routing is true.

C representation: XBERGOcrRotation is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrRotation does not appear anywhere in the generated header.

Rotation information for an OCR element.

Field Type Default Description
angle_degrees double — Rotation angle in degrees (0, 90, 180, 270 for PaddleOCR).
confidence double* NULL Confidence score for the rotation detection.

C representation: XBERGOcrTable is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrTable does not appear anywhere in the generated header.

Table detected via OCR.

Represents a table structure recognized during OCR processing.

Field Type Default Description
cells const char* — Table cells as a 2D vector (rows × columns)
markdown const char* — Markdown representation of the table
page_number uint32_t — Page number where the table was found (1-indexed)
bounding_box XBERGAlefHandle /* serde(default) */ Bounding box of the table in pixel coordinates (from OCR word positions).

C representation: XBERGOcrTableBoundingBox is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOcrTableBoundingBox does not appear anywhere in the generated header.

Bounding box for an OCR-detected table in pixel coordinates.

Field Type Default Description
left uint32_t — Left x-coordinate (pixels)
top uint32_t — Top y-coordinate (pixels)
right uint32_t — Right x-coordinate (pixels)
bottom uint32_t — Bottom y-coordinate (pixels)

C representation: XBERGOrientationResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGOrientationResult does not appear anywhere in the generated header.

Document orientation detection result.

Field Type Default Description
degrees uint32_t — Detected orientation in degrees (0, 90, 180, or 270).
confidence float — Confidence score (0.0-1.0).

C representation: XBERGPaddleOcrConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPaddleOcrConfig does not appear anywhere in the generated header.

Configuration for PaddleOCR backend.

Configures PaddleOCR text detection and recognition with multi-language support. Uses a builder pattern for convenient configuration.

Field Type Default Description
language const char* "en" Language code (e.g., “en”, “ch”, “jpn”, “kor”, “deu”, “fra”)
cache_dir const char* NULL Optional Hugging Face Hub cache root for model files. When unset, the standard HF_HUB_CACHE, legacy HUGGINGFACE_HUB_CACHE, and HF_HOME conventions are used.
use_angle_cls int32_t false Enable angle classification for rotated text (default: false). Can misfire on short text regions, rotating crops incorrectly before recognition.
enable_table_detection int32_t false Enable table structure detection (default: false, unlike TesseractConfig which defaults to true — see crate.ocr.types.TesseractConfig). PaddleOCR has no per-word table-candidate confidence carve-out the way Tesseract’s TSV does (every recognised word is a clustering candidate), so unfiltered clustering over-fabricates tables on ordinary prose. This stays off until it clears a cell-scored accuracy measurement (see the pinning test at the bottom of this file). Consequence for callers: switching an existing config from the Tesseract backend to PaddleOCR with stock defaults silently produces no OCR tables — enable this explicitly if you need them. When enabled, this only clusters and reconstructs a grid from word boxes PaddleOCR has already produced (see crate.paddle_ocr.backend.PaddleOcrBackend.process_image) — it does not run any additional model inference, so the added per-page cost is the CPU-only clustering/reconstruction pass, not another ONNX call.
det_db_thresh float 0.3 Database threshold for text detection (default: 0.3) Range: 0.0-1.0, higher values require more confident detections
det_db_box_thresh float 0.5 Box threshold for text bounding box refinement (default: 0.5) Range: 0.0-1.0
det_db_unclip_ratio float 1.6 Unclip ratio for expanding text bounding boxes (default: 1.6) Controls the expansion of detected text regions
det_limit_side_len uint32_t 1024 Maximum side length for detection image (default: 1024) Larger images may be resized to this limit for faster inference
rec_batch_num uint32_t 6 Batch size for recognition inference (default: 6) Number of text regions to process simultaneously
padding uint32_t 10 Padding in pixels added around the image before detection (default: 10). Large values can include surrounding content like table gridlines.
drop_score float 0.5 Minimum recognition confidence score for text lines (default: 0.5). Text regions with recognition confidence below this threshold are discarded. Matches PaddleOCR Python’s drop_score parameter. Range: 0.0-1.0
model_tier const char* "mobile" Model tier controlling detection/recognition model size and accuracy trade-off. For PP-OCRv5 (model_version = "pp-ocrv5"): - "mobile" (default): Lightweight models (~4.5MB detection, ~16.5MB recognition), fast download and inference - "server": Large, high-accuracy models (~88MB detection, ~84MB recognition), best for GPU or complex documents For PP-OCRv6 (model_version = "pp-ocrv6", the default): - "small": ~9.9MB detection, full 18,708-char CJK+Latin+JA/KO recognition dictionary. The default "mobile" resolves here, so this is what an unconfigured extraction uses. - "medium": ~62MB detection, same dictionary. Higher accuracy, substantially slower on CPU. A legacy "server" tier, or any unrecognised value, resolves here. - "tiny": ~1.8MB detection, but a reduced 6,904-char (~zh/en) dictionary — it cannot read the scripts the other two cover. Note PaddleOCR pages do not run concurrently: the ONNX session is held behind a mutex, so the thread budget goes to intra-op parallelism and wall time scales with page count times per-page inference. Tier choice therefore dominates throughput on multi-page documents.
model_version const char* "pp-ocrv6" Model generation: "pp-ocrv6" (default) or "pp-ocrv5". PP-OCRv6 adds a unified CJK+Latin+JA/KO recognition model with medium/small/tiny tiers (see model_tier). Scripts outside the v6 unified coverage (Arabic, Cyrillic, Devanagari, Greek, Tamil, Telugu, Thai) transparently fall back to the PP-OCRv5 per-script recognition models. Defaults to "pp-ocrv6"; the default model_tier ("mobile") resolves to the v6 "small" tier. Select "pp-ocrv5" to pin the legacy per-script/unified fleet.
inference_backend XBERGAlefHandle NULL Explicit inference engine choice. NULL (the default) resolves to the compiled default: ort when the paddle-ocr-ort feature is compiled in, otherwise tract. An explicit choice is validated against the compiled features when the OCR engine is constructed (see crate.paddle_ocr.backend.effective_backend); requesting an engine whose feature is not compiled in is a clear configuration error rather than a silent fallback.

C representation: XBERGPageBoundary is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageBoundary does not appear anywhere in the generated header.

Byte offset boundary for a page.

Tracks where a specific page’s content starts and ends in the main content string, enabling mapping from byte positions to page numbers. Offsets are guaranteed to be at valid UTF-8 character boundaries when using standard String methods (push_str, push, etc.).

Field Type Default Description
byte_start uintptr_t — Byte offset where this page starts in the content string (UTF-8 valid boundary, inclusive)
byte_end uintptr_t — Byte offset where this page ends in the content string (UTF-8 valid boundary, exclusive)
page_number uint32_t — Page number (1-indexed)

C representation: XBERGPageClassification is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageClassification does not appear anywhere in the generated header.

Classification result for a single page.

Field Type Default Description
page_number uint32_t — 1-indexed page number this classification belongs to.
labels const char* — Labels assigned to the page. Single-label classification yields exactly one entry; multi-label classification yields any subset of the configured label set.

Since: v1.0

C representation: XBERGPageClassificationConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageClassificationConfig does not appear anywhere in the generated header.

Configuration for the page-classification post-processor.

Field Type Default Description
prompt_template const char* NULL Minijinja prompt template. Receives {{ labels }} (joined list), {{ page_text }} and {{ multi_label }} variables. NULL lets the backend pick a sensible default.
labels const char* — The set of labels the classifier may emit. Must contain at least one entry.
multi_label int32_t /* serde(default) */ Allow multiple labels per page. Single-label mode returns at most one label.
llm XBERGAlefHandle — LLM configuration used for classification.

C representation: XBERGPageConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageConfig does not appear anywhere in the generated header.

Page extraction and tracking configuration.

Controls how pages are extracted, tracked, and represented in the extraction results. When NULL, page tracking is disabled.

Page range tracking in chunk metadata (first_page/last_page) is automatically enabled when page boundaries are available and chunking is configured.

Field Type Default Description
extract_pages int32_t false Extract pages as separate array (ExtractedDocument.pages)
insert_page_markers int32_t false Insert page markers in main content string
marker_format const char* "<!-- PAGE {page_num} -->" Page marker format (use {page_num} placeholder) Default: “\n\n\n\n”

C representation: XBERGPageContent is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageContent does not appear anywhere in the generated header.

Content for a single page/slide.

When page extraction is enabled, documents are split into per-page content with associated tables and images mapped to each page.

Uses shared tables and images for memory efficiency:

  • const Table* enables zero-copy sharing of table data
  • const ExtractedImage* enables zero-copy sharing of image data
  • Maintains exact JSON compatibility via custom Serialize/Deserialize

This reduces memory overhead for documents with shared tables/images by avoiding redundant copies during serialization.

Field Type Default Description
page_number uint32_t — Page number (1-indexed)
content const char* — Text content for this page
tables const char* /* serde(default) */ Tables found on this page (uses Arc for memory efficiency) Serializes as const Table* for JSON compatibility while maintaining shared in-memory ownership for zero-copy sharing.
image_indices const char* /* serde(default) */ Indices into ExtractedDocument.images for images found on this page. Each value is a zero-based index into the top-level images collection. Only populated when extract_images = true in the extraction config.
image_preprocessing XBERGAlefHandle /* serde(default) */ OCR image preprocessing applied to this page’s raster.
hierarchy XBERGAlefHandle NULL Hierarchy information for the page (when hierarchy extraction is enabled) Contains text hierarchy levels (H1-H6) extracted from the page content.
is_blank int32_t* NULL Whether this page is blank (no meaningful text content) Determined during extraction based on text content analysis. A page is blank if it has fewer than 3 non-whitespace characters and contains no tables or images.
layout_regions const char* NULL Layout detection regions for this page (when layout detection is enabled). Contains detected layout regions with class, confidence, bounding box, and area fraction. Only populated when layout detection is configured.
speaker_notes const char* NULL Speaker notes for this slide (PPTX only). Contains the text from the slide’s notes pane (ppt/notesSlides/notesSlide{N}.xml). Only populated when the source is a PPTX file and notes are present.
section_name const char* NULL Section name this slide belongs to (PPTX only). PowerPoint sections group slides into logical chapters (<p:sectionLst> in ppt/presentation.xml). Only populated when the source is a PPTX file and the slide belongs to a named section.
sheet_name const char* NULL Sheet name for this page (XLSX/ODS only). Each spreadsheet sheet maps to one PageContent entry. This field carries the sheet’s display name as it appears in the workbook. NULL for all non-spreadsheet formats and for sheets with an empty name.
ocr_confidence XBERGAlefHandle /* serde(default) */ Aggregate OCR confidence for this page. NULL when the page was not OCR’d.

C representation: XBERGPageDimensions is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageDimensions does not appear anywhere in the generated header.

Metadata for individual page/slide/sheet.

Captures per-page information including dimensions, content counts, and visibility state (for presentations).

Field Type Default Description
width double — Page width in points or pixels.
height double — Page height in points or pixels.

C representation: XBERGPageHierarchy is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageHierarchy does not appear anywhere in the generated header.

Page hierarchy structure containing heading levels and block information.

Used when PDF text hierarchy extraction is enabled. Contains hierarchical blocks with heading levels (H1-H6) for semantic document structure.

Field Type Default Description
block_count uint32_t — Number of hierarchy blocks on this page
blocks const char* /* serde(default) */ Hierarchical blocks with heading levels

C representation: XBERGPageInfo is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageInfo does not appear anywhere in the generated header.

Field Type Default Description
number uint32_t — Page number (1-indexed)
title const char* NULL Page title (usually for presentations)
dimensions XBERGAlefHandle /* serde(default) */ Dimensions in points (PDF) or pixels (images).
image_count uint32_t* NULL Number of images on this page
table_count uint32_t* NULL Number of tables on this page
hidden int32_t* NULL Whether this page is hidden (e.g., in presentations)
is_blank int32_t* NULL Whether this page is blank (no meaningful text, no images, no tables) A page is considered blank if it has fewer than 3 non-whitespace characters and contains no tables or images. This is useful for filtering out empty pages in scanned documents or PDFs with blank separator pages.
has_vector_graphics int32_t /* serde(default) */ Whether this page contains non-trivial vector graphics (paths, shapes, curves) Indicates the presence of vector-drawn content such as charts, diagrams, or geometric shapes (e.g., from Adobe InDesign, LaTeX TikZ). These are invisible to ExtractedDocument.images since they are not embedded as raster XObjects. Set to true when path count exceeds a heuristic threshold, signaling that downstream consumers may want to rasterize the page to capture this content. Only populated for PDFs; NULL for other document types.

C representation: XBERGPageOcrConfidence is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageOcrConfidence does not appear anywhere in the generated header.

Aggregate OCR legibility score for a page, reported by the backend that produced its text.

This is distinct from OcrConfidence, which scores a single detected element (a word or line) using detection/recognition confidence from the OCR engine itself. PageOcrConfidence is a page-level summary computed after noise filtering, intended for triage of which pages are worth a closer look, not for comparing OCR engines against each other.

Field Type Default Description
score double* NULL Aggregate legibility score in 0.0..=1.0, or NULL when the backend that produced this page does not report a calibrated legibility scale. Backends differ in what their confidence numbers mean (see OcrConfidence’s per-backend normalization), and not every backend maps onto a 0.0-1.0 legibility scale at all. When a backend has no such calibrated scale, this is NULL rather than a misleading number, and scores must never be compared across backends.
word_count uint32_t — Number of words the score was averaged over, AFTER noise filtering. A small word_count means the average is based on little evidence, so a high score next to a small word_count is not representative of the whole page.
backend const char* — Name of the OCR backend that produced the page text.

C representation: XBERGPageRange is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageRange does not appear anywhere in the generated header.

Page range for a chunk (0-indexed, inclusive).

Field Type Default Description
start uint32_t — Start page (0-indexed, inclusive).
end uint32_t — End page (0-indexed, inclusive).

C representation: XBERGPageSignals is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageSignals does not appear anywhere in the generated header.

Per-page signals extracted from PDF content.

Field Type Default Description
page_number uint32_t — 1-indexed page number.
text_excerpt const char* — First ~500 characters of extracted text.
starts_with_letterhead_like int32_t — true if page starts with letterhead-like content (ALL CAPS line in first 5 lines or a logo-image bbox at top).
has_page_number_one_marker int32_t — true if text contains “Page 1” or “1 of N” pattern.
has_signature_block int32_t — true if text contains signature indicators (“Sincerely”, “Signed”) or a signature image bbox.
layout_text_density float — Text density: characters per page area, normalised to [0.0, 1.0].

C representation: XBERGPageSpan is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageSpan does not appear anywhere in the generated header.

A single page covered by a chunk, with an optional bounding box on that page.

See ChunkMetadata.page_spans (#1295) for population semantics.

Field Type Default Description
page uint32_t — Page number (1-indexed).
bbox XBERGAlefHandle NULL Bounding box on this page, if known.

C representation: XBERGPageStructure is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPageStructure does not appear anywhere in the generated header.

Unified page structure for documents.

Supports different page types (PDF pages, PPTX slides, Excel sheets) with character offset boundaries for chunk-to-page mapping.

Field Type Default Description
total_count uint32_t — Total number of pages/slides/sheets
unit_type XBERGAlefHandle — Type of paginated unit
boundaries const char* NULL Character offset boundaries for each page Maps character ranges in the extracted content to page numbers. Used for chunk page range calculation.
pages const char* NULL Detailed per-page metadata (optional, only when needed)

C representation: XBERGPatternMatch is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPatternMatch does not appear anywhere in the generated header.

One detected PII span in the input text.

Field Type Default Description
start uintptr_t — Inclusive byte-offset start of the match in the source text.
end uintptr_t — Exclusive byte-offset end of the match.
category XBERGAlefHandle — Category the match belongs to.
text const char* — Matched substring (owned copy — pattern engine returns owned data so the caller can free the original text if needed before replacement).

C representation: XBERGPdfAnnotation is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPdfAnnotation does not appear anywhere in the generated header.

A PDF annotation extracted from a document page.

Field Type Default Description
annotation_type XBERGAlefHandle — The type of annotation.
content const char* NULL Text content of the annotation (e.g., comment text, link URL).
page_number uint32_t — Page number where the annotation appears (1-indexed).
bounding_box XBERGAlefHandle NULL Bounding box of the annotation on the page.
author const char* /* serde(default) */ Author/creator of the annotation (PDF /T entry).
modified const char* /* serde(default) */ Last modification date of the annotation (PDF /M entry), as a raw PDF date string (e.g. "D:20240115120000Z").
color const char* /* serde(default) */ Annotation colour (PDF /C entry), normalised to a CSS-compatible #rrggbb hex string. Gray and CMYK colour spaces are converted to RGB.
subject const char* /* serde(default) */ Subject of the annotation (PDF /Subj entry).
quad_points const char* /* serde(default) */ Per-line bounding boxes derived from the annotation’s /QuadPoints entry. Present for text markup annotations (Highlight, Underline, StrikeOut, Squiggly), one box per marked line/run of text.
marked_text const char* /* serde(default) */ The document text covered by Self.quad_points, recovered from the page content underneath the marked-up region. Populated for Highlight, Underline, StrikeOut, and Squiggly annotations when the underlying text could be recovered.

C representation: XBERGPdfConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPdfConfig does not appear anywhere in the generated header.

PDF-specific configuration.

Field Type Default Description
extract_images int32_t false Extract images from PDF
extract_tables int32_t true Extract tables from PDF. When true (default), runs the native engine’s grid detector and, if it finds nothing, falls back to the heuristic text-layer reconstruction in pdf.native.table.extract_tables_heuristic. Set to false to skip both passes — tables will then be empty in the result.
passwords const char* NULL List of passwords to try when opening encrypted PDFs
extract_metadata int32_t true Extract PDF metadata
hierarchy XBERGAlefHandle NULL Hierarchy extraction configuration (None = hierarchy extraction disabled)
extract_annotations int32_t false Extract PDF annotations (text notes, highlights, links, stamps). Default: false
top_margin_fraction float* NULL Top margin fraction (0.0–1.0) of page height to exclude headers/running heads. Ignored when ContentFilterConfig.include_headers is true. Effective nonzero margins require per-page OCR so geometry can be filtered; document-capable OCR backends use their image-processing path in that case. Default: 0.0 (disabled; set explicitly to filter header content)
bottom_margin_fraction float* NULL Bottom margin fraction (0.0–1.0) of page height to exclude footers/page numbers. Ignored when ContentFilterConfig.include_footers is true. Effective nonzero margins require per-page OCR so geometry can be filtered; document-capable OCR backends use their image-processing path in that case. Default: 0.0 (disabled; set explicitly to filter footer content)
allow_single_column_tables int32_t false Allow single-column pseudo tables in extraction results. By default, tables with fewer than 2 columns (layout-guided) or 3 columns (heuristic) are rejected. When true, the minimum column count is relaxed to 1, allowing single-column structured data (glossaries, itemized lists) to be emitted as tables. Other quality filters (density, sparsity, prose detection) still apply.
ocr_inline_images int32_t false Perform OCR on inline images extracted from PDF pages and attach the recognized text to each ExtractedImage.ocr_result. Uses the backend selected by ExtractionConfig.ocr, or the default OCR backend when no OCR configuration is supplied. Requires the ocr or ocr-pipeline feature. Per-image failures degrade gracefully (the image is returned without OCR text rather than failing the whole extraction). Default: false.
extract_form_fields int32_t true Extract AcroForm and XFA form fields into ExtractedDocument.form_fields. When true (default), reads the document’s interactive form structure (field names, types, values, widget geometry). Cheap and strictly additive — non-form PDFs simply yield an empty list. Set to false to skip the form pass entirely.
reading_order int32_t false Reorder extracted text by layout-detected reading order. When true, projects text spans onto layout-detected regions, performs column detection, and emits spans in natural reading order (important for multi-column academic PDFs). It also repairs 90/180/270-degree rotated text runs — sideways tables and captions — that otherwise read word-reversed and glued (GH#1358); see crate.extractors.pdf.reading_order for the rotation-handling details and its limits. Requires the layout-detection feature and a page for which layout detection actually produces hints: a page with no detected regions falls back to the original, unrepaired extraction order even with this enabled. Independent of LayoutStrategy, which only controls whether layout detection runs at all — enabling Always or Auto alone does not turn reordering on. Defaults to false.
backend XBERGAlefHandle XBERG_NATIVE Which engine parses and renders this PDF. Defaults to PdfBackend.Native. Selecting PdfBackend.Pdfium requires the pdf-pdfium feature and is rejected otherwise; the pdfium engine is also deliberately narrower in scope than Native – see PdfBackend and extractors.pdf.pdfium_engine for details.

C representation: XBERGPdfFormField is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPdfFormField does not appear anywhere in the generated header.

A form field extracted from a PDF’s AcroForm or XFA structure.

Populated by the PDF extractor when PdfConfig.extract_form_fields is enabled and the document is a fillable form. Supports both AcroForm (standard) and XFA (XML Forms Architecture) layers. When both are present, AcroForm fields take priority (canonical fallback per PDF spec), and XFA-only fields are appended. The collection is empty for non-form PDFs and for non-PDF formats.

Field Type Default Description
name const char* — Partial field name (the leaf name within the field hierarchy).
full_name const char* — Fully-qualified field name (dotted path from the form root).
field_type XBERGAlefHandle — Classified field type.
value const char* /* serde(default) */ Current field value, if any.
default_value const char* /* serde(default) */ Default field value, if any.
flags uint32_t /* serde(default) */ Raw field-flags bitmask (read-only, required, multiline, …).
page uint32_t* /* serde(default) */ 1-indexed page the field’s widget appears on. Currently always NULL for AcroForm fields; page assignment is a deferred enhancement requiring spatial analysis of widget annotations per page.
bbox XBERGAlefHandle /* serde(default) */ Widget bounding box on its page, if known.
max_length uint32_t* /* serde(default) */ Maximum input length for text fields, if specified.
tooltip const char* /* serde(default) */ Tooltip / alternate field description, if present.

C representation: XBERGPdfMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPdfMetadata does not appear anywhere in the generated header.

PDF-specific metadata.

Contains metadata fields specific to PDF documents that are not in the common Metadata structure. Common fields like title, authors, keywords, and dates are at the Metadata level.

Field Type Default Description
pdf_version const char* NULL PDF version (e.g., “1.7”, “2.0”)
producer const char* NULL PDF producer (application that created the PDF)
is_encrypted int32_t* NULL Whether the PDF is encrypted/password-protected
width int64_t* NULL First page width in points (1/72 inch)
height int64_t* NULL First page height in points (1/72 inch)
page_count uint32_t* NULL Total number of pages in the PDF document
scanned_confidence float* NULL How strongly the document’s most scan-like page resembles a scan, in [0.0, 1.0]. NULL when the document could not be inspected. A full-page raster with no visible text scores at least 0.85; a born-digital slide with a full-bleed background image scores 0.50.
scanned_pages const char* NULL Pages that look like scans (1-indexed), using the default confidence threshold. NULL when the document could not be inspected; empty when no page qualifies.
fabricated_text_pages const char* NULL Pages whose text was dominated by fabricated character mappings (1-indexed): MappingProvenance.Fallback, a font whose glyph-to-Unicode mapping resolved to a value the extractor chose rather than read from the file (issue #1254). This is a fact about how the text was derived, independent of scanned_pages’s raster-based scan detection, and independent of whether the resulting text happens to look structurally like prose (issue #1667: a broken mapping that lands on ordinary letters and punctuation passes every character-shape check but is still fabricated). NULL when OcrQualityThresholds.enable_provenance_ocr_routing is false; empty when no page qualifies.
implausible_text_pages const char* NULL Pages whose native text layer reads as no real detectable language (1-indexed): a /ToUnicode CMap (or other mapping tier) that resolves every glyph to a character, consistently the WRONG one (e.g. a ROT-shifted mapping), so the page is structurally indistinguishable from real prose to fabricated_text_pages’s provenance check and to every character-shape heuristic (issue #1696; issue #1667’s quality_score: 1.0 with no warning on such a page is the same underlying gap). This is a content-plausibility fact, independent of fabricated_text_pages and of scanned_pages. NULL when OcrQualityThresholds.enable_plausibility_ocr_routing is false; empty when no page qualifies.
layout_gated_pages const char* NULL Pages the auto layout strategy skipped (1-indexed). NULL unless layout detection ran with LayoutStrategy.Auto; empty when the gate selected every page.
layout_gate_reasons const char* NULL Why the auto layout gate selected or skipped each page. Index i is page i + 1. Snake_case values such as multi_column, table_grid, or plain_text (the skip reason). NULL unless layout detection ran with LayoutStrategy.Auto.

C representation: XBERGPixelDimensions is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPixelDimensions does not appear anywhere in the generated header.

Image preprocessing metadata.

Tracks the transformations applied to an image during OCR preprocessing, including DPI normalization, resizing, and resampling.

Field Type Default Description
width uintptr_t — Width in pixels.
height uintptr_t — Height in pixels.

C representation: XBERGPlugin is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPlugin does not appear anywhere in the generated header.

Base trait that all plugins must implement.

This trait provides common functionality for plugin lifecycle management, identification, and metadata.

All plugins must be Send + Sync to support concurrent usage across threads.

Returns the unique name/identifier for this plugin.

The name should be:

  • Unique across all plugins
  • Lowercase with hyphens (e.g., “my-custom-plugin”)
  • URL-safe characters only

Signature:

const char* xberg_plugin_name(XBERGAlefHandle this);

Example:

const char *result = xberg_plugin_name(instance);

Returns: const char*

Returns the semantic version of this plugin.

Should follow semver format: MAJOR.MINOR.PATCH

Defaults to the xberg crate version.

Signature:

const char* xberg_plugin_version(XBERGAlefHandle this);

Example:

const char *result = xberg_plugin_version(instance);

Returns: const char*

Initialize the plugin.

Called once when the plugin is registered. Use this to:

  • Load configuration
  • Initialize resources (connections, caches, etc.)
  • Validate dependencies

This method takes &self instead of &mut self to work with Arc<dyn Plugin>. Plugins needing mutable state during initialization should use interior mutability patterns (Mutex, RwLock, OnceCell, etc.).

Errors:

Should return an error if initialization fails. The plugin will not be registered if this method returns an error.

Defaults to a no-op for stateless plugins.

Signature:

int32_t xberg_plugin_initialize(XBERGAlefHandle this);

Example:

xberg_plugin_initialize(instance);

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.

Shutdown the plugin.

Called when the plugin is being unregistered or the application is shutting down. Use this to:

  • Close connections
  • Flush caches
  • Release resources

This method takes &self instead of &mut self to work with Arc<dyn Plugin>. Plugins needing mutable state during shutdown should use interior mutability patterns (Mutex, RwLock, etc.).

Errors:

Errors during shutdown are logged but don’t prevent the shutdown process.

Defaults to a no-op for stateless plugins.

Signature:

int32_t xberg_plugin_shutdown(XBERGAlefHandle this);

Example:

xberg_plugin_shutdown(instance);

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.

Optional plugin description for debugging and logging.

Defaults to empty string if not overridden.

Signature:

const char* xberg_plugin_description(XBERGAlefHandle this);

Example:

const char *result = xberg_plugin_description(instance);

Returns: const char*

Optional plugin author information.

Defaults to empty string if not overridden.

Signature:

const char* xberg_plugin_author(XBERGAlefHandle this);

Example:

const char *result = xberg_plugin_author(instance);

Returns: const char*


C representation: XBERGPostProcessor is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPostProcessor does not appear anywhere in the generated header.

Trait for post-processor plugins.

Post-processors transform or enrich extraction results after the initial extraction is complete. They can:

  • Clean and normalize text
  • Add metadata (language, keywords, entities)
  • Split content into chunks
  • Score quality
  • Apply custom transformations

Post-processors are executed in stage order:

  1. Early - Language detection, entity extraction
  2. Middle - Keyword extraction, token reduction
  3. Late - Custom hooks, final validation

Within each stage, processors are executed in registration order.

Post-processor errors are non-fatal by default - they’re captured in metadata and execution continues. To make errors fatal, return an error from process().

Post-processors must be thread-safe (Send + Sync).

Process an extraction result.

Transform or enrich the extraction result. Can modify:

  • content - The extracted text
  • metadata - Add or update metadata fields
  • tables - Modify or enhance table data

Returns:

Ok(()) if processing succeeded, Err(...) for fatal failures.

Errors:

Return errors for fatal processing failures. Non-fatal errors should be captured in metadata directly on the result.

This signature avoids unnecessary cloning of large extraction results by taking a mutable reference instead of ownership. Processors modify the result in place.

Signature:

int32_t xberg_post_processor_process(XBERGAlefHandle this, XBERGAlefHandle result, XBERGAlefHandle config);

Example:

xberg_post_processor_process(instance, 0, 0);

Parameters:

Name Type Required Description
result XBERGAlefHandle Yes Mutable reference to the extraction result to process
config XBERGAlefHandle Yes Extraction configuration

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.

Get the processing stage for this post-processor.

Determines when this processor runs in the pipeline.

Returns:

The ProcessingStage (Early, Middle, or Late).

Signature:

XBERGAlefHandle xberg_post_processor_processing_stage(XBERGAlefHandle this);

Example:

XBERGAlefHandle result = xberg_post_processor_processing_stage(instance);

Returns: XBERGAlefHandle

Optional: Check if this processor should run for a given result.

Allows conditional processing based on MIME type, metadata, or content. Defaults to true (always run).

Returns:

true if the processor should run, false to skip.

Signature:

int32_t xberg_post_processor_should_process(XBERGAlefHandle this, XBERGAlefHandle result, XBERGAlefHandle config);

Example:

int32_t result = xberg_post_processor_should_process(instance, 0, 0);

Parameters:

Name Type Required Description
result XBERGAlefHandle Yes The extracted document
config XBERGAlefHandle Yes The extraction config

Returns: int32_t

xberg_post_processor_estimated_duration_ms()
Section titled “xberg_post_processor_estimated_duration_ms()”

Optional: Estimate processing time in milliseconds.

Used for logging and debugging. Defaults to 0 (unknown).

Returns:

Estimated processing time in milliseconds.

Signature:

uint64_t xberg_post_processor_estimated_duration_ms(XBERGAlefHandle this, XBERGAlefHandle result);

Example:

uint64_t result = xberg_post_processor_estimated_duration_ms(instance, 0);

Parameters:

Name Type Required Description
result XBERGAlefHandle Yes The extracted document

Returns: uint64_t

Execution priority within the processing stage.

Higher values run first within the same ProcessingStage. Defaults to 50. Use 0-49 for fallback processors, 50 for normal processors, and 51-255 for high-priority processors that should run early in their stage.

Signature:

int32_t xberg_post_processor_priority(XBERGAlefHandle this);

Example:

int32_t result = xberg_post_processor_priority(instance);

Returns: int32_t


C representation: XBERGPostProcessorConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPostProcessorConfig does not appear anywhere in the generated header.

Post-processor configuration.

Field Type Default Description
enabled int32_t true Enable post-processors
enabled_processors const char* NULL Whitelist of processor names to run (None = all enabled)
disabled_processors const char* NULL Blacklist of processor names to skip (None = none disabled)
enabled_set const char* NULL Pre-computed AHashSet for O(1) enabled processor lookup
disabled_set const char* NULL Pre-computed AHashSet for O(1) disabled processor lookup

C representation: XBERGPptxAppProperties is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPptxAppProperties does not appear anywhere in the generated header.

Application properties from docProps/app.xml for PPTX

Contains PowerPoint-specific document metadata.

Field Type Default Description
application const char* NULL Application name (e.g., “Microsoft Office PowerPoint”)
app_version const char* NULL Application version
total_time int32_t* NULL Total editing time in minutes
company const char* NULL Company name
doc_security int32_t* NULL Document security level
scale_crop int32_t* NULL Scale crop flag
links_up_to_date int32_t* NULL Links up to date flag
shared_doc int32_t* NULL Shared document flag
hyperlinks_changed int32_t* NULL Hyperlinks changed flag
slides int32_t* NULL Number of slides
notes int32_t* NULL Number of notes
hidden_slides int32_t* NULL Number of hidden slides
multimedia_clips int32_t* NULL Number of multimedia clips
presentation_format const char* NULL Presentation format (e.g., “Widescreen”, “Standard”)
slide_titles const char* NULL Slide titles

C representation: XBERGPptxExtractionResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPptxExtractionResult does not appear anywhere in the generated header.

PowerPoint (PPTX) extraction result.

Contains extracted slide content, metadata, and embedded images/tables.

Field Type Default Description
content const char* — Extracted text content from all slides
metadata XBERGAlefHandle — Presentation metadata
slide_count uintptr_t — Total number of slides
image_count uintptr_t — Total number of embedded images
table_count uintptr_t — Total number of tables
images const char* — Extracted images from the presentation
page_structure XBERGAlefHandle NULL Slide structure with boundaries (when page tracking is enabled)
page_contents const char* NULL Per-slide content (when page tracking is enabled)
document XBERGAlefHandle NULL Structured document representation
hyperlinks const char* /* serde(default) */ Hyperlinks discovered in slides.
office_metadata const char* /* serde(default) */ Office metadata extracted from docProps/core.xml and docProps/app.xml. Contains keys like “title”, “author”, “created_by”, “subject”, “keywords”, “modified_by”, “created_at”, “modified_at”, etc.
revisions const char* /* serde(default) */ Slide comments as revisions. Each <p:cm> element in ppt/comments/comment{N}.xml becomes a DocumentRevision { kind: Comment } with author (resolved from ppt/commentAuthors.xml), ISO-8601 timestamp, and RevisionAnchor.Slide { index }. NULL when no comment XML parts exist.

C representation: XBERGPptxMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPptxMetadata does not appear anywhere in the generated header.

PowerPoint presentation metadata.

Extracted from PPTX files containing slide counts and presentation details.

Field Type Default Description
slide_count uint32_t — Total number of slides in the presentation
slide_names const char* NULL Names of slides (if available)
image_count uint32_t* NULL Number of embedded images
table_count uint32_t* NULL Number of tables

C representation: XBERGPreprocessingOptions is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPreprocessingOptions does not appear anywhere in the generated header.

HTML preprocessing options for document cleanup before conversion.

Field Type Default Description
enabled int32_t true Enable HTML preprocessing globally
preset XBERGAlefHandle XBERG_STANDARD Preprocessing preset level (Minimal, Standard, Aggressive)
remove_navigation int32_t true Remove navigation elements (nav, breadcrumbs, menus, sidebars)
remove_forms int32_t true Remove form elements (forms, inputs, buttons, etc.)

C representation: XBERGPresentationHyperlink is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPresentationHyperlink does not appear anywhere in the generated header.

A hyperlink discovered in a presentation slide.

Field Type Default Description
url const char* — Link destination.
label const char* NULL Optional visible label.

C representation: XBERGPreset is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPreset does not appear anywhere in the generated header.

A curated structured-extraction preset loaded from the embedded library.

Each preset is a JSON file under src/presets/library/<id>/v1.json that validates against the meta-schema in src/presets/preset.schema.json.

Downstream catalog consumers can inject presets via extend_from_dir. The embedded OSS library ships only the generic_document toy preset.

Field Type Default Description
id const char* — Stable, URL-safe preset identifier (lowercase snake_case).
version const char* — Monotonic version string (e.g. v1).
schema_name const char* — Human-readable schema name forwarded to the LLM as the response/tool name.
description const char* — One-line preset description shown in the registry UI.
category XBERGAlefHandle — Top-level category for grouping in the playground.
tags const char* /* serde(default) */ Free-form tags used for search/filtering. May be empty.
schema const char* — JSON Schema (Draft 2020-12) describing the structured output shape.
system_prompt const char* — Instruction primer sent to the model.
context_template const char* /* serde(default) */ Optional mustache-style template merged with caller-supplied context.
merge_mode XBERGAlefHandle — Strategy for merging per-batch outputs across paginated calls.
preferred_call_mode XBERGAlefHandle — Default call mode suggested for this preset; heuristics may override.
emit_citations int32_t — When true, the prompt asks the model to wrap each field as {value, page, bbox, confidence} for downstream citation overlays.
sample XBERGAlefHandle /* serde(default) */ Optional bundled sample (input file + reference output) for preview.
fingerprint const char* /* serde(default) */ Stable sha256 fingerprint of the canonical preset file contents. Populated at registry load — not present in the on-disk JSON files. Used as a cache-invalidation token by the worker pipeline.

C representation: XBERGPresetSample is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPresetSample does not appear anywhere in the generated header.

Pointer to a sample input + its reference output bundled with the preset.

Field Type Default Description
input_path const char* — Path to the sample input file, relative to the preset directory.
output_path const char* — Path to the reference structured output, relative to the preset directory.

C representation: XBERGPresetSummary is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPresetSummary does not appear anywhere in the generated header.

Lightweight projection of Preset used by the registry list endpoint (omits the full schema and prompt to keep the payload small).

Field Type Default Description
id const char* — Preset identifier matching Preset.id.
version const char* — Preset version matching Preset.version.
schema_name const char* — Schema name matching Preset.schema_name.
description const char* — One-line preset description.
category XBERGAlefHandle — Top-level category.
tags const char* — Free-form tags.
preferred_call_mode XBERGAlefHandle — Default call mode.
emit_citations int32_t — Whether the preset prompts the model for citations.
fingerprint const char* — Stable fingerprint matching Preset.fingerprint.

C representation: XBERGProcessingWarning is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGProcessingWarning does not appear anywhere in the generated header.

A non-fatal warning from a processing pipeline stage.

Captures errors from optional features that don’t prevent extraction but may indicate degraded or incomplete results. Inspect these independently from ExtractedDocument.quality_score, which assesses retained text only.

Field Type Default Description
source const char* — The pipeline stage or feature that produced this warning (e.g., “embedding”, “chunking”, “language_detection”, “output_format”).
message const char* — Human-readable description of what went wrong.

C representation: XBERGPropertyChange is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPropertyChange does not appear anywhere in the generated header.

A single run-level or style-level property change.

Used for revisions that change formatting rather than text content. from and to store normalized property values when the source format exposes them; either side may be absent when the format only records one side of the change.

Field Type Default Description
name const char* — Property name, such as "bold", "italic", "font_size", or "font_color".
from const char* NULL Value before the change, when available.
to const char* NULL Value after the change, when available.

C representation: XBERGProxyConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGProxyConfig does not appear anywhere in the generated header.

Proxy configuration for HTTP requests.

Field Type Default Description
url const char* — Proxy URL (e.g. “http://proxy:8080”, “socks5://proxy:1080”).
username const char* NULL Optional username for proxy authentication.
password const char* NULL Optional password for proxy authentication.

C representation: XBERGPstMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGPstMetadata does not appear anywhere in the generated header.

Outlook PST archive metadata.

Field Type Default Description
message_count uintptr_t — Total number of email messages found in the PST archive.

C representation: XBERGQrBoundingBox is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGQrBoundingBox does not appear anywhere in the generated header.

Pixel-space bounding box of a QR code inside its source image.

Field Type Default Description
x uint32_t — Horizontal pixel offset of the bounding box top-left corner.
y uint32_t — Vertical pixel offset of the bounding box top-left corner.
width uint32_t — Width of the bounding box in pixels.
height uint32_t — Height of the bounding box in pixels.

C representation: XBERGQrCode is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGQrCode does not appear anywhere in the generated header.

One QR code decoded from an extracted image.

Field Type Default Description
payload const char* — Decoded payload (text, URL, vCard string, …).
confidence float* NULL Detector-reported confidence in [0.0, 1.0]. NULL when the decoder does not expose confidence (the default rqrr backend always reports Some because successful decode implies high confidence).
bbox XBERGAlefHandle NULL Bounding box of the QR code inside the source image, in pixel coordinates (x, y of the top-left corner; width, height of the rectangle). NULL if the decoder did not report a bounding box.

C representation: XBERGRakeParams is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRakeParams does not appear anywhere in the generated header.

RAKE-specific parameters.

Field Type Default Description
min_word_length uintptr_t 1 Minimum word length to consider (default: 1).
max_words_per_phrase uintptr_t 3 Maximum words in a keyword phrase (default: 3).

C representation: XBERGRecognizedTable is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRecognizedTable does not appear anywhere in the generated header.

Pre-computed table markdown for a table detection region.

Produced by the TATR-based table structure recognizer and surfaced as part of layout-aware OCR results. The struct lives here (under layout-types, pure-Rust) so that consumers who do not enable layout-detection (ORT) can still reference the type in their own code.

Field Type Default Description
detection_bbox XBERGAlefHandle — Detection bbox that this table corresponds to (for matching).
cells const char* — Table cells as a 2D vector (rows × columns).
markdown const char* — Rendered markdown table.

Since: v1.0

C representation: XBERGRedactionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRedactionConfig does not appear anywhere in the generated header.

Configuration for the redaction post-processor.

Field Type Default Description
categories const char* NULL Categories to redact. Empty means “every category supported by the engine.”
strategy XBERGAlefHandle XBERG_MASK Strategy applied to every match.
ner XBERGAlefHandle NULL Optional NER backend — required to redact PERSON / ORGANIZATION / LOCATION categories (the pure-Rust pattern engine only covers regex-detectable PII).
preserve_offsets int32_t true When true, chunk byte ranges are kept consistent with the rewritten content by adjusting byte_start / byte_end after replacement. When false, chunk byte ranges still refer to the original content offsets — useful when downstream consumers want to map findings back to the original document.
custom_terms const char* NULL Arbitrary user-supplied literal terms to redact. Each term is treated as a regex hit against the document, surfacing as PiiCategory.Custom(label) in RedactionFinding where label is the per-term label (defaulting to the literal value itself). Case-insensitive by default; set RedactionTerm.case_sensitive for exact match. Use this when you need to redact tenant-specific tokens (employee IDs, project codes, internal product names) without writing a custom plugin.
custom_patterns const char* NULL Arbitrary user-supplied regex patterns to redact. Same surfacing semantics as custom_terms: each hit becomes a PiiCategory.Custom(label) finding. Patterns are validated at config-construction time via RedactionConfig.validate.

C representation: XBERGRedactionFinding is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRedactionFinding does not appear anywhere in the generated header.

One redaction event: which span was rewritten, why, and with what.

Field Type Default Description
start uint32_t — Byte-offset start in the original (pre-redaction) ExtractedDocument.content.
end uint32_t — Byte-offset end (exclusive) in the original ExtractedDocument.content.
category XBERGAlefHandle — PII category that fired this redaction.
strategy XBERGAlefHandle — Strategy applied to this finding (mask, hash, token-replace, drop).
replacement_token const char* — String that replaced the original mention. Always present; for Drop the replacement is the empty string.

C representation: XBERGRedactionPattern is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRedactionPattern does not appear anywhere in the generated header.

One user-supplied regex pattern to redact.

The pattern is compiled with the Rust regex crate (no look-around). Case sensitivity is encoded in the pattern via the (?i) inline flag when Self.case_sensitive is false.

Field Type Default Description
label const char* — Custom category label surfaced in RedactionFinding.category.
pattern const char* — Regex pattern (Rust regex crate dialect — no look-around).
case_sensitive int32_t false When true, match case-sensitively; otherwise prepend (?i) to the regex.

C representation: XBERGRedactionReport is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRedactionReport does not appear anywhere in the generated header.

Audit report describing what the redaction processor found and how it replaced it.

The redactor returns this alongside the rewritten content so compliance, replay, and audit-log consumers can see exactly what fired. Offsets are relative to the original pre-redaction content and are intended for audit reconstruction only — the original bytes are dropped at the end of the pipeline.

Field Type Default Description
findings const char* — Individual redaction findings in original-source byte order.
total_redacted uint32_t — Total number of redactions applied across the document.

C representation: XBERGRedactionTerm is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRedactionTerm does not appear anywhere in the generated header.

One user-supplied literal term to redact.

Matched as a regex-escaped substring (so callers do not need to escape metacharacters themselves). Case-insensitive by default — set Self.case_sensitive to true for exact byte-match semantics.

Field Type Default Description
label const char* — Custom category label surfaced in RedactionFinding.category.
value const char* — Literal value to match. Regex metacharacters are escaped automatically.
case_sensitive int32_t false When true, match the value as-is; otherwise match ASCII-case-insensitively.

C representation: XBERGRegistry is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRegistry does not appear anywhere in the generated header.

Sorted map of preset id → Preset.

Build the registry from preset files embedded at compile time under src/presets/library/. Validates every file against the meta-schema.

Signature:

XBERGAlefHandle xberg_registry_load_embedded();

Example:

XBERGAlefHandle result = xberg_registry_load_embedded();

Returns: XBERGAlefHandle

Errors: Returns the sentinel handle 0 on error.

Return the global registry, loading it on first access.

Panics:

Panics if any embedded preset is malformed. The build-time validation test ensures this cannot happen for the embedded presets; a panic here indicates a build artifact problem, not a runtime error.

Signature:

XBERGAlefHandle xberg_registry_global();

Example:

XBERGAlefHandle result = xberg_registry_global();

Returns: XBERGAlefHandle

Look up a preset by its identifier.

Signature:

XBERGAlefHandle xberg_registry_get(XBERGAlefHandle this, const char* id);

Example:

XBERGAlefHandle result = xberg_registry_get(instance, "value");

Parameters:

Name Type Required Description
id const char* Yes The id

Returns: XBERGAlefHandle

Materialize a PresetSummary list for the public registry endpoint.

Signature:

const char* xberg_registry_summaries(XBERGAlefHandle this);

Example:

const char* result = xberg_registry_summaries(instance);

Returns: const char*

Number of presets currently loaded.

Signature:

uintptr_t xberg_registry_len(XBERGAlefHandle this);

Example:

uintptr_t result = xberg_registry_len(instance);

Returns: uintptr_t

Whether the registry contains zero presets.

Signature:

int32_t xberg_registry_is_empty(XBERGAlefHandle this);

Example:

int32_t result = xberg_registry_is_empty(instance);

Returns: int32_t

Read raw sample bytes for <preset_id> from library/<id>/samples/<name>. Returns NULL when the file is absent.

Signature:

const uint8_t* xberg_registry_sample_bytes(XBERGAlefHandle this, const char* preset_id, const char* name);

Example:

const uint8_t* result = xberg_registry_sample_bytes(instance, "value", "value");

Parameters:

Name Type Required Description
preset_id const char* Yes The preset id
name const char* Yes The name

Returns: const uint8_t*

Load additional preset files from a runtime directory and insert them into this registry.

Reads every *.json file directly under dir (non-recursive), validates each against the meta-schema, and inserts it. Files that fail validation are rejected — the error is returned immediately and the registry is left in a partially-updated state. Existing entries with the same id are overwritten.

Returns the number of presets successfully loaded from dir.

This is the injection point for downstream catalogs that add curated presets on top of the single embedded OSS preset.

Signature:

uintptr_t xberg_registry_extend_from_dir(XBERGAlefHandle this, const char* dir);

Example:

uintptr_t result = xberg_registry_extend_from_dir(instance, "value");

Parameters:

Name Type Required Description
dir const char* Yes The dir

Returns: uintptr_t

Errors: Returns 0 on error.


C representation: XBERGRenderer is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRenderer does not appear anywhere in the generated header.

Trait for document renderers that convert extraction results to output strings.

Renderers are typically stateless converters that transform extracted content into a specific output format (Markdown, HTML, Djot, plain text, etc.). They participate in the standard Plugin lifecycle so custom renderers can be registered from any supported binding language.

The format name is exposed via Plugin.name. For stateless renderers the Plugin lifecycle methods (version, initialize, shutdown) all take no-op defaults and need not be overridden.

Renderers must be Send + Sync (inherited from Plugin).

Binding-safe rendering entry point for foreign-language plugin bridges.

Accepts one public extraction result and returns the rendered output.

Signature:

const char* xberg_renderer_render_result(XBERGAlefHandle this, XBERGAlefHandle result);

Example:

const char *result = xberg_renderer_render_result(instance, 0);

Parameters:

Name Type Required Description
result XBERGAlefHandle Yes The extracted document

Returns: const char*

Errors: Returns NULL on error.


C representation: XBERGRerankedDocument is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRerankedDocument does not appear anywhere in the generated header.

A single document returned by the reranker, with its position in the input and score.

index maps back to the caller’s original document list, so metadata arrays (e.g. IDs, paths) can be reordered without passing them through the reranker.

Field Type Default Description
index uintptr_t — Position of this document in the original input documents slice.
score float — Relevance score in [0, 1]. Higher means more relevant to the query.
document const char* — The document text.

C representation: XBERGRerankerBackend is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRerankerBackend does not appear anywhere in the generated header.

Trait for in-process reranker backend plugins.

Cross-encoders score (query, document) pairs jointly and return a raw logit per document. The crate-level rerank dispatcher applies sigmoid to convert logits to [0, 1] scores, sorts descending by score, and truncates to top_k.

Async to match the convention used by EmbeddingBackend and other plugin traits. Host-language bridges wrap their synchronous host callables in spawn_blocking or the equivalent.

Backends must be Send + Sync + 'static. They are stored in Arc<dyn RerankerBackend> and may be called concurrently from xberg’s dispatcher. If the backend’s underlying model is not thread-safe, the backend itself must serialize access internally (e.g. via Mutex<Inner>).

  • rerank(query, documents) MUST return exactly documents.len() scores. The dispatcher validates this before sorting and returning to callers; a non-conforming backend surfaces as a XbergError.Validation, not a panic.

  • Scores are raw logits in any range — callers must NOT assume [0, 1]. The dispatcher applies sigmoid before sorting.

  • rerank may be called from any thread. Its future must be Send (enforced by async_trait when #[async_trait] is used on non-WASM targets).

  • shutdown() (inherited from Plugin) may be invoked concurrently with an in-flight rerank() call. Implementations must tolerate this — letting in-flight calls finish via the Arc reference and only releasing shared state that isn’t needed by rerank.

The synchronous rerank entry uses tokio.task.block_in_place to await the trait’s async rerank, which requires a multi-thread tokio runtime. Callers running inside a current_thread runtime must use rerank_async instead.

Score a list of documents against a query.

Returns one raw logit per document in the same order as the input. The dispatcher applies sigmoid to convert to [0, 1] scores.

Errors:

Implementations should return Plugin for backend-specific failures. The dispatcher validates the returned length against documents.len() before sorting.

Signature:

const char* xberg_reranker_backend_rerank(XBERGAlefHandle this, const char* query, const char* documents);

Example:

const char* result = xberg_reranker_backend_rerank(instance, "value", NULL);

Parameters:

Name Type Required Description
query const char* Yes The query
documents const char* Yes The documents

Returns: const char*

Errors: Returns NULL on error.


C representation: XBERGRerankerConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRerankerConfig does not appear anywhere in the generated header.

Configuration for the reranking pipeline.

Controls which model to use, how many results to return, and download/cache behavior for local ONNX models.

Field Type Default Description
model XBERGAlefHandle XBERG_PRESET { name: "balanced" } The reranker model to use (defaults to “balanced” preset if not specified).
top_k uintptr_t* NULL Return at most this many documents. NULL returns all. Applied after sorting by score, so the highest-scoring documents are kept.
batch_size uintptr_t 32 Batch size for local ONNX cross-encoder inference.
show_download_progress int32_t false Show model download progress (local ONNX path only). When enabled, transfer progress for the model, tokenizer and config files is reported at info level on the xberg.model_download target while they download (#279). A warm Hugging Face cache transfers nothing and so reports nothing. Ignored by RerankerModelType.Llm and RerankerModelType.Plugin, which download no model.
cache_dir const char* NULL Optional alternate Hugging Face cache root for model files. When unset, hf-hub follows the standard Hugging Face environment and platform cache conventions.
acceleration XBERGAlefHandle NULL Hardware acceleration for the reranker ONNX model. Controls which execution provider (CPU, CUDA, CoreML, TensorRT) is used for local inference. Defaults to NULL (auto-select per platform).
max_rerank_duration_secs uint64_t* 60 Maximum wall-clock duration (in seconds) for a single rerank() call when using RerankerModelType.Plugin. Applies only to the in-process plugin path — protects against hung host-language backends. On timeout, the dispatcher returns Plugin instead of blocking forever. NULL disables the timeout. The default (60 seconds) is conservative for common in-process inference; increase for large document sets on slow hardware.

C representation: XBERGResolvedPreset is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGResolvedPreset does not appear anywhere in the generated header.

A preset merged with caller-supplied overrides (custom schema, prompt suffix, context map). Output is what the pipeline orchestrator consumes.

Field Type Default Description
id const char* — Source preset identifier.
version const char* — Source preset version.
fingerprint const char* — Fingerprint of the source preset file, used as a cache token.
schema_name const char* — Schema name forwarded to the LLM.
schema const char* — Effective JSON Schema (caller override or the preset’s own).
system_prompt const char* — System prompt with rendered context appended.
merge_mode XBERGAlefHandle — Merge strategy for paginated outputs.
preferred_call_mode XBERGAlefHandle — Preferred call mode.
emit_citations int32_t — Whether the prompt asks for per-field citations.

C representation: XBERGRevisionDelta is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGRevisionDelta does not appear anywhere in the generated header.

The content changes that make up a single revision.

For insertions and deletions the content field carries the added/removed lines as DiffLine.Added / DiffLine.Removed entries. For format changes, property_changes carries normalized before/after formatting values when the source document exposes them.

Field Type Default Description
content const char* NULL Line-level content changes for this revision.
table_changes const char* NULL Cell-level table changes for this revision.
property_changes const char* NULL Formatting or metadata property changes for this revision.

C representation: XBERGSecurityLimits is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSecurityLimits does not appear anywhere in the generated header.

Configuration for security limits across extractors.

All limits are intentionally conservative to prevent DoS attacks while still supporting legitimate documents.

Field Type Default Description
max_archive_size uintptr_t 524288000 Maximum uncompressed size for archives (500 MB)
max_compression_ratio uintptr_t 100 Maximum compression ratio before flagging as potential bomb (100:1)
max_files_in_archive uintptr_t 10000 Maximum number of files in archive (10,000)
max_nesting_depth uintptr_t 1024 Maximum nesting depth for structures (1024)
max_entity_length uintptr_t 1048576 Maximum length of any single XML entity / attribute / token (1 MiB). This is a per-token cap, NOT a total cap — billion-laughs class attacks where a single entity expands to hundreds of MB are caught here, while normal long text content (a paragraph, a CDATA block) is caught by max_content_size instead.
max_content_size uintptr_t 104857600 Maximum string growth and decoded image allocation per operation (100 MB). Per-page passes such as layout detection charge each batch against this limit, not the whole document; only max_pages bounds the rasters retained across a document (GH#1721).
max_iterations uintptr_t 10000000 Maximum iterations per operation
max_xml_depth uintptr_t 1024 Maximum XML depth (1024 levels)
max_table_cells uintptr_t 100000 Maximum aggregate table cells per document (100,000). Raise this for trusted large tabular inputs. Higher values permit proportionally more parsing work and output allocation.
max_pages uintptr_t* NULL Maximum number of pages (or slides, or frames) in a single document. NULL means unlimited. Checked once the count is known and before any per-page work (OCR, layout detection, rendering) starts. Byte-size limits do not bound page count: a scanned page can compress to a few kilobytes, so a document well under max_content_size or max_archive_size can still hold thousands of pages of per-page work. Defaults to NULL (unlimited) because a real ceiling here is workload-specific and a low default would silently reject legitimate large documents; callers that want a ceiling set this explicitly. Enforced for: PDF (extractors.pdf, page count via xberg_native_pdf/lopdf), PPTX (extraction.pptx, slide count from the archive’s slide parts), Keynote (extractors.iwork.keynote, slide count from Index/Slide-*.iwa entry names), ODP (extractors.odp, draw:page count in content.xml), and multi-frame TIFF images built with the ocr feature (extractors.image, frame count via the tiff crate). Not enforced for any other format, including DOCX, ODT, XLSX, legacy PPT/DOC, Pages/Numbers, and TIFF images when the ocr feature is disabled: those formats either have no fixed “page” the crate can count without doing the expensive work itself (DOCX/ODT page count is a layout outcome, not a stored value), or have no per-page pipeline to gate at all. Setting max_pages on a document of an unenforced format is silently a no-op, not a guarantee.

C representation: XBERGServerConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGServerConfig does not appear anywhere in the generated header.

API server configuration.

This struct holds all configuration options for the Xberg API server, including host/port settings, CORS configuration, and upload limits.

  • host: “127.0.0.1” (localhost only)
  • port: 8000
  • cors_origins: empty vector (allows all origins)
  • max_request_body_bytes: 104_857_600 (100 MB)
  • max_multipart_field_bytes: 104_857_600 (100 MB)
  • job_timeout_secs: 600 (10 minutes)
Field Type Default Description
host const char* "127.0.0.1" Server host address (e.g., “127.0.0.1”, “0.0.0.0”)
port uint16_t 8000 Server port number
cors_origins const char* NULL CORS allowed origins. Empty vector means allow all origins. If this is an empty vector, the server will accept requests from any origin. If populated with specific origins (e.g., "<https://example.com>"), only those origins will be allowed.
max_request_body_bytes uintptr_t 104857600 Maximum size of request body in bytes (default: 100 MB)
max_multipart_field_bytes uintptr_t 104857600 Maximum size of multipart fields in bytes (default: 100 MB)
job_timeout_secs uint64_t 600 Fallback timeout, in seconds, for POST /extract-async jobs whose request does not pin down extraction_timeout_secs (default: 600, 10 minutes). A per-request extraction_timeout_secs: Some(n) always overrides this value. An explicit extraction_timeout_secs: null deliberately does NOT mean “run unbounded” — it still falls back to this server-configured cap, because an unbounded job on a shared server is a denial-of-service risk.

C representation: XBERGSitemapUrl is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSitemapUrl does not appear anywhere in the generated header.

A URL entry from a sitemap.

Field Type Default Description
url const char* — The URL.
lastmod const char* NULL The last modification date, if present.
changefreq const char* NULL The change frequency, if present.
priority const char* NULL The priority, if present.

C representation: XBERGSparseEmbedding is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSparseEmbedding does not appear anywhere in the generated header.

A sparse learned embedding: vocabulary term indices and their weights.

indices are ascending vocabulary token ids; values[i] is the weight for indices[i]. The two arrays always have equal length. Only strictly-positive terms are retained, so the representation is genuinely sparse.

Field Type Default Description
indices const char* — Vocabulary token ids with non-zero weight, ascending.
values const char* — Weights parallel to SparseEmbedding.indices.

C representation: XBERGSparseEmbeddingConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSparseEmbeddingConfig does not appear anywhere in the generated header.

Configuration for the sparse-embedding pipeline.

Controls which model to use, batching, and download/cache behavior for the local ONNX SPLADE model.

Field Type Default Description
model XBERGAlefHandle XBERG_PRESET The sparse-embedding model to use (defaults to the “opensearch-v3-distill” preset).
batch_size uintptr_t 16 Batch size for local ONNX inference. SPLADE emits a [seq, vocab] logit tensor per document, so memory scales with batch size — keep this modest.
max_length uintptr_t 256 Maximum token sequence length for the tokenizer.
show_download_progress int32_t false Show model download progress (local ONNX path only). When enabled, transfer progress for the model, tokenizer and config files is reported at info level on the xberg.model_download target while they download (#279). A warm Hugging Face cache transfers nothing and so reports nothing. Ignored by SparseEmbeddingModelType.Plugin, which downloads no model.
cache_dir const char* NULL Optional alternate Hugging Face cache root for model files. When unset, hf-hub follows the standard Hugging Face environment and platform cache conventions.
acceleration XBERGAlefHandle NULL Hardware acceleration for the sparse-embedding ONNX model.
max_embed_duration_secs uint64_t* 60 Maximum wall-clock duration (in seconds) for a single embed call when using SparseEmbeddingModelType.Plugin. NULL disables the timeout.

C representation: XBERGSparseEmbeddingPreset is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSparseEmbeddingPreset does not appear anywhere in the generated header.

Static metadata for a bundled SPLADE preset (WASM/Android-safe, no ORT).

Field Type Default Description
name const char* — Stable preset name referenced from config.
model_repo const char* — HuggingFace repository hosting the ONNX model.
model_file const char* — Path to the ONNX file within the repo.
additional_files const char* — Sibling files that must be downloaded alongside model_file.
max_length uintptr_t — Maximum token sequence length.
description const char* — Human-readable description.

C representation: XBERGSsrfPolicy is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSsrfPolicy does not appear anywhere in the generated header.

SSRF policy configuration.

Field Type Default Description
deny_private int32_t true If true, reject URLs that resolve to private/metadata IP ranges.
allowlist const char* NULL Hostnames and IP ranges permitted regardless of deny_private. The allowlist is an override of deny_private, not an intersection with it. Precedence, in order: 1. deny_private == false permits everything; the allowlist is not consulted. 2. A hostname matching an Exact or Suffix entry is permitted immediately, before DNS resolution — so the deny-list is never applied to it. This trusts the host string: a name that resolves into private space is still permitted. 3. A literal or resolved IP inside a Cidr entry is permitted even though it is in the default deny-list. 4. Otherwise the default deny-list decides. An empty allowlist therefore denies nothing by itself — it simply leaves deny_private and the deny-list in sole control.
max_redirects uint8_t 5 Maximum number of HTTP redirects to follow during validation.
scheme_allowlist const char* ["http", "https"] Allowed URI schemes. Default: ["http", "https"]. Only http and https are supported. An empty list denies every URL.

C representation: XBERGStructuredData is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGStructuredData does not appear anywhere in the generated header.

Structured data (Schema.org, microdata, RDFa) block.

Field Type Default Description
data_type XBERGAlefHandle — Type of structured data
raw_json const char* — Raw JSON string representation
schema_type const char* NULL Schema type if detectable (e.g., “Article”, “Event”, “Product”)

C representation: XBERGStructuredDataResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGStructuredDataResult does not appear anywhere in the generated header.

Result of parsing a structured data file (JSON, JSONL, YAML, or TOML).

Field Type Default Description
content const char* — The extracted text content, formatted for readability.
format const char* — The source format identifier (e.g. "json", "yaml", "toml").
metadata const char* — Key-value metadata extracted from recognized text fields.
text_fields const char* — JSON paths of fields that were classified as text-bearing.
value const char* NULL The parsed document as a canonical serde_json.Value tree, when the source format could be represented as one. NULL only for TOML inputs whose toml.Value fails to round-trip through serde_json.Value (xberg-io/xberg#155): the extractor falls back to a raw code block in that case.
flattened const char* — Flattened path: value renderings for every leaf field, in traversal order. Previously computed and discarded (xberg-io/xberg#166); now surfaced so callers get a full-text view even when the structured renderer only emits headings/lists for a subset of fields.

C representation: XBERGStructuredExtractionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGStructuredExtractionConfig does not appear anywhere in the generated header.

Configuration for LLM-based structured data extraction.

Sends extracted document content to a VLM with a JSON schema, returning structured data that conforms to the schema.

Field Type Default Description
schema const char* — JSON Schema defining the desired output structure.
schema_name const char* "extraction" Schema name passed to the LLM’s structured output mode.
schema_description const char* /* serde(default) */ Optional schema description for the LLM.
strict int32_t /* serde(default) */ Enable strict mode — output must exactly match the schema.
prompt const char* /* serde(default) */ Custom Jinja2 extraction prompt template. When NULL, a default template is used. Available template variables: - {{ content }} — The extracted document text. - {{ schema }} — The JSON schema as a formatted string. - {{ schema_name }} — The schema name. - {{ schema_description }} — The schema description (may be empty).
llm XBERGAlefHandle — LLM configuration for the extraction.

Since: v1.0

C representation: XBERGSummarizationConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSummarizationConfig does not appear anywhere in the generated header.

Configuration for the summarisation post-processor.

Field Type Default Description
strategy XBERGAlefHandle XBERG_EXTRACTIVE Summarisation strategy.
max_tokens uint32_t* NULL Maximum summary length in tokens. NULL lets the backend pick a default.
llm XBERGAlefHandle NULL LLM configuration for the abstractive backend. Ignored when strategy = Extractive. Required when strategy = Abstractive.

C representation: XBERGSupportedFormat is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSupportedFormat does not appear anywhere in the generated header.

A supported document format entry.

Represents a file extension and its corresponding MIME type that Xberg can process.

Field Type Default Description
extension const char* — File extension (without leading dot), e.g., “pdf”, “docx”
mime_type const char* — MIME type string, e.g., “application/pdf”

C representation: XBERGSvgOptions is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGSvgOptions does not appear anywhere in the generated header.

SVG-specific configuration for the image-encode pipeline.

Applies when the source image is SVG or when the output format is set to ImageOutputFormat.Svg. Available when the svg feature is active.

Used via ImageExtractionConfig.svg.

Field Type Default Description
sanitize int32_t true Run SVG bytes through usvg sanitization (strips external href attributes, JavaScript event handlers, and foreignObject elements) even when the output format is Native. Defaults to true.
render_dpi float 96 Target DPI when rasterizing SVG to a pixel-based format (PNG, JPEG, WebP, HEIF). The tree’s viewBox is scaled by render_dpi / 96.0 before the pixel buffer is allocated. Defaults to 96.0 (1× CSS pixel density).

C representation: XBERGTable is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTable does not appear anywhere in the generated header.

Extracted table structure.

Represents a table detected and extracted from a document (PDF, image, etc.). Tables are converted to both structured cell data and Markdown format.

Field Type Default Description
cells const char* NULL Table cells as a 2D vector (rows × columns)
markdown const char* — Markdown representation of the table
page_number uint32_t — Page number where the table was found (1-indexed)
bounding_box XBERGAlefHandle NULL Bounding box of the table’s position. Only populated when position data is available from the producing extractor. The coordinate space depends on how the table was produced, and callers must know which route produced a given Table before interpreting this field: - Tables extracted from a PDF’s native content, and tables recognized on a scanned PDF page that went through xberg’s OCR pipeline (--force-ocr / --ocr-scanned-pages), are in PDF points with a bottom-left origin (x0=left, y0=bottom, x1=right, y1=top; y increases upward). For the OCR case, the pipeline rescales the backend’s raw pixel output into this space before it reaches Table.bounding_box — see rescale_ocr_bboxes_to_page_points in extractors.pdf.ocr. - Tables detected by OCR on a standalone image with no backing PDF page (for example extracting a bare PNG/JPEG/TIFF) are in raster pixel coordinates with a top-left origin (x0=left, y0=top, x1=right, y1=bottom; y increases downward) — the same convention the OCR backend (Tesseract, PaddleOCR, candle-based backends) or the layout detector reported them in. There is no PDF page geometry to rescale into for this case, so the raw pixel box is passed through unchanged.
table_id const char* NULL Stable identifier shared by every tables[] entry that represents a fragment of the same physical table. Assigned deterministically by the extraction pipeline (e.g. a sequential "table-N" in document order); never derived from randomness or wall-clock time, so the same input document always produces the same ids. Consumers can use it to reconcile the markdown blocks in content / pages[].content / chunks[].content with the structured entries in tables[]. NULL when the extractor did not assign one. Today, same-page fragments of one physical table are already merged into a single tables[] entry before ids are assigned (see PDF table stitching), so in practice table_id is unique per entry rather than shared across several. A table split across a page boundary is intentionally not linked — its per-page pieces get separate ids. Sharing one id across page-boundary fragments is a known possible future extension, not implemented yet.
cell_styles const char* NULL Paragraph styles carried by individual cells, for the cells that have one. Sparse and flat on purpose. A DOCX banner row – row 0, one cell spanning the grid, styled Heading1..Heading6 – is what Word’s navigation pane and a TOC field treat as the document outline, but as a table cell it reached consumers as anonymous text (GH#1587). cells keeps the bare text: prefixing it with # would put a markdown heading inside a table cell, which is invalid where it lands and would change text every existing consumer already reads. This list is the signal instead. Entries are only emitted for cells that actually carry a style, so an ordinary table serialises exactly as it did before. Indices are into cells.
columns const char* NULL Header cells for this fragment, i.e. the first row of cells. Populated even when this fragment’s own header row was merged away or physically lives in a sibling fragment (see table_id), so a single fragment is interpretable in isolation. NULL when no header row could be determined.

C representation: XBERGTableCell is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTableCell does not appear anywhere in the generated header.

Individual table cell with content and optional styling.

Future extension point for rich table support with cell-level metadata.

Field Type Default Description
content const char* — Cell content as text
row_span uint32_t 1 Row span (number of rows this cell spans)
col_span uint32_t 1 Column span (number of columns this cell spans)
is_header int32_t — Whether this is a header cell

C representation: XBERGTableCellStyle is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTableCellStyle does not appear anywhere in the generated header.

The paragraph style a single table cell’s text carries, located by grid position.

Flat rather than a nested const Vec<Option<..*>>: the nested shape marshals badly across the FFI bindings, and the data is sparse anyway. See Table.cell_styles.

Field Type Default Description
row uint32_t — Zero-indexed row of the cell this style belongs to.
col uint32_t — Zero-indexed column of the cell this style belongs to.
heading_level uint8_t* NULL Outline level 1-6 when the style resolves to a heading, otherwise NULL.
style_name const char* NULL Human-readable style name, e.g. heading 2.

C representation: XBERGTableDiff is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTableDiff does not appear anywhere in the generated header.

Cell-level changes for a pair of tables that share the same index.

Field Type Default Description
from_index uintptr_t — Zero-based index of the table in both a.tables and b.tables.
to_index uintptr_t — Zero-based index in b.tables (equal to from_index for same-dimension tables).
cell_changes const char* — Cell-level changes within the table.

C representation: XBERGTableGrid is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTableGrid does not appear anywhere in the generated header.

Structured table grid with cell-level metadata.

Stores row/column dimensions and a flat list of cells with position info.

Field Type Default Description
rows uint32_t — Number of rows in the table.
cols uint32_t — Number of columns in the table.
cells const char* NULL All cells in row-major order.

C representation: XBERGTesseractConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTesseractConfig does not appear anywhere in the generated header.

Tesseract OCR configuration.

Provides fine-grained control over Tesseract OCR engine parameters. Most users can use the defaults, but these settings allow optimization for specific document types (invoices, handwriting, etc.).

This is the public-facing counterpart of ocr.types.TesseractConfig (the internal, engine-facing representation with u8/String fields instead of i32/const char**). They are two independent struct definitions bridged only by an explicit From<&TesseractConfig> for crate.ocr.types.TesseractConfig impl in ocr/types.rs — that conversion carries field values across, but each struct keeps its own Default impl, and the conversion does nothing to keep those two defaults in sync. Several production call sites (extractors.image.apply_default_tesseract_psm, configured_region_ocr, sparse_image_ocr_fallback_config) construct this struct’s default and convert it, bypassing the internal struct’s Default entirely — so if the two defaults disagree, standalone image OCR silently uses this struct’s value while PDF-embedded OCR (which can reach the internal Default directly when no tesseract_config is set) uses the other. When changing a default here, also update ocr.types.TesseractConfig.default, and vice versa.

Field Type Default Description
language const char* ["eng"] Language code(s) for OCR recognition. For Tesseract, languages are joined with “+”. A list is the canonical form and the only form accepted by the binding object APIs (Python, Node, PHP, WASM, etc.): ["eng", "deu"]. When deserializing from a config file, JSON body, or the REST/MCP API, a single string is also accepted, either as one code (“eng”) or “+”-joined (“eng+deu”).
psm int32_t* NULL Page Segmentation Mode (1-13). PSM 0 is rejected: Tesseract’s PSM_OSD_ONLY performs orientation and script detection with no character recognition, so it cannot satisfy a text-extraction request and previously yielded an empty document (GH#1586). NULL (the default) means the caller made no explicit choice: the extraction pipeline applies its own context-appropriate PSM (whole-image PSM 11, vertical- language PSM 5, layout-region PSM 6, or the sparse-text retry’s PSM 3) exactly as it would with no TesseractConfig at all — see issue #1573. Setting any other field on this struct no longer changes that behaviour. A rendered PDF page (force_ocr / force_ocr_pages / scanned-page OCR) that is one full-page scan gets the same whole-image PSM as standalone image OCR of that raster (11, or 5 for a vertical language), so the scan needs no explicit psm (GH#1786). A page that is not a scan, such as forced OCR of a vector page, keeps the engine’s automatic-layout PSM (3). Common explicit values: - 3: Fully automatic page segmentation (native engine default) - 6: Assume a single uniform block of text (WASM engine default — avoids layout-analysis hang) - 11: Sparse text with no particular order
output_format const char* "markdown" Output format (“text” or “markdown”)
oem int32_t 3 OCR Engine Mode (0-3). - 0: Legacy engine only - 1: Neural nets (LSTM) only (usually best) - 2: Legacy + LSTM - 3: Default (based on what’s available)
min_confidence double 0 Minimum confidence threshold (0.0-100.0). Words with confidence below this threshold may be rejected or flagged.
preprocessing XBERGAlefHandle NULL Image preprocessing configuration. Controls how images are preprocessed before OCR. Can significantly improve quality for scanned documents or low-quality images.
enable_table_detection int32_t true Enable automatic table detection and reconstruction
table_min_confidence double 0 Minimum confidence threshold for table detection (0.0-1.0)
table_column_threshold int32_t 50 Column threshold for table detection (pixels)
table_row_threshold_ratio double 0.5 Row threshold ratio for table detection (0.0-1.0)
use_cache int32_t true Enable OCR result caching
classify_use_pre_adapted_templates int32_t true Use pre-adapted templates for character classification
language_model_ngram_on int32_t true Enable N-gram language model. Kept on by default (see Self.default and ocr.types.TesseractConfig.language_model_ngram_on for the rationale); keep this field’s default in sync with the internal struct’s.
tessedit_dont_blkrej_good_wds int32_t true Don’t reject good words during block-level processing
tessedit_dont_rowrej_good_wds int32_t true Don’t reject good words during row-level processing
tessedit_enable_dict_correction int32_t true Enable dictionary correction
tessedit_char_whitelist const char* "" Whitelist of allowed characters (empty = all allowed)
tessedit_char_blacklist const char* "" Blacklist of forbidden characters (empty = none forbidden)
tessedit_use_primary_params_model int32_t true Use primary language params model
textord_space_size_is_variable int32_t true Variable-width space detection
thresholding_method int32_t 0 Tesseract image-binarization method (0-2): 0 = Otsu (default), 1 = LeptonicaOtsu, 2 = Sauvola. Sent to Tesseract’s integer thresholding_method engine variable. GH#1784: used to be a bool sent as "true"/"false", silently ignored by Tesseract’s integer parser. A config for the old field still deserializes: false -> 0 (Otsu, the prior no-op), true -> 1 (LeptonicaOtsu, the old “adaptive” doc).

C representation: XBERGTextAnnotation is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTextAnnotation does not appear anywhere in the generated header.

Inline text annotation — byte-range based formatting and links.

Annotations reference byte offsets into the node’s text content, enabling precise identification of formatted regions.

Field Type Default Description
start uint32_t — Start byte offset in the node’s text content (inclusive).
end uint32_t — End byte offset in the node’s text content (exclusive).
kind XBERGAlefHandle — Annotation type.

C representation: XBERGTextExtractionResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTextExtractionResult does not appear anywhere in the generated header.

Plain text and Markdown extraction result.

Contains the extracted text along with statistics and, for Markdown files, structural elements like headers and links.

Field Type Default Description
content const char* — Extracted text content
line_count uintptr_t — Number of lines
word_count uintptr_t — Number of words
character_count uintptr_t — Number of characters
headers const char* NULL Markdown headers (text only, Markdown files only)
links const char* NULL Markdown links (Markdown files only).
code_blocks const char* NULL Code blocks (Markdown files only).

C representation: XBERGTextMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTextMetadata does not appear anywhere in the generated header.

Text/Markdown metadata.

Extracted from plain text and Markdown files. Includes word counts and, for Markdown, structural elements like headers and links.

Field Type Default Description
line_count uint32_t — Number of lines in the document
word_count uint32_t — Number of words
character_count uint32_t — Number of characters
headers const char* NULL Markdown headers (headings text only, for Markdown files)
links const char* NULL Markdown links (for Markdown files).
code_blocks const char* NULL Code blocks (for Markdown files).

C representation: XBERGTokenCounter is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTokenCounter does not appear anywhere in the generated header.

Per-category running counter for RedactionStrategy.TokenReplace.

Create a fresh counter with no previous state.

Signature:

XBERGAlefHandle xberg_token_counter_new();

Example:

XBERGAlefHandle result = xberg_token_counter_new();

Returns: XBERGAlefHandle


C representation: XBERGTokenReductionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTokenReductionConfig does not appear anywhere in the generated header.

Configuration for the token-reduction pipeline.

Field Type Default Description
level XBERGAlefHandle XBERG_MODERATE Reduction intensity level.
language_hint const char* NULL ISO 639-1 language code hint for stopword selection (e.g. "en", "de").
preserve_markdown int32_t false Preserve Markdown formatting tokens during reduction.
preserve_code int32_t true Preserve code block contents unchanged.
semantic_threshold float 0.3 Cosine similarity threshold below which sentences are considered dissimilar.
enable_parallel int32_t true Use Rayon parallel iterators for multi-core processing.
use_simd int32_t true Use SIMD-optimized text scanning where available.
custom_stopwords const char* NULL Per-language custom stopword lists (language_code → stopword_list).
preserve_patterns const char* NULL Regex patterns whose matched text is always preserved unchanged.
target_reduction float* NULL Target fraction of text to retain (0.0–1.0); NULL = no fixed target.
enable_semantic_clustering int32_t false Group semantically similar sentences and emit only one per cluster.
preserve_important_words int32_t true Skip removal of words with “important” characteristics (all-caps acronyms, words containing digits, mixed-case identifiers, very long words) during the Aggressive/Maximum common-word removal pass. true (the default) protects those words even when they would otherwise be dropped as low-value filler; false lets the frequency/ length heuristics apply uniformly to every word, including ones that look like acronyms or technical terms (#269).

C representation: XBERGTokenReductionOptions is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTokenReductionOptions does not appear anywhere in the generated header.

Token reduction configuration.

Field Type Default Description
mode const char* "off" Reduction mode: “off”, “light”, “moderate”, “aggressive”, “maximum”
preserve_important_words int32_t true Preserve important words (capitalized, technical terms)

C representation: XBERGTokenizerBackend is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTokenizerBackend does not appear anywhere in the generated header.

Trait for in-process tokenizer backend plugins.

Unlike EmbeddingBackend, this trait is synchronous: the chunk splitter calls Self.count_tokens inside its boundary search, many times per chunk, so counting must be a direct call with no async dispatch. Host-language bridges (PyO3, napi-rs, etc.) invoke their host callable synchronously on the calling thread; implementations should keep count_tokens cheap — it dominates chunking time when the backend is slow.

initialize() is called once during registration, before any count_tokens call; lazy-loading implementations should load their vocabulary there. After registration succeeds, count_tokens may be called from any thread, concurrently. shutdown() runs on unregistration and may overlap an in-flight count_tokens call from a chunking run that resolved the backend earlier — implementations must tolerate this, e.g. by keeping the resources count_tokens needs alive via Arc.

Backends must be Send + Sync + 'static (inherited from Plugin). They are stored in Arc<dyn TokenizerBackend> and called concurrently from xberg’s chunking pipeline. If the underlying tokenizer isn’t thread-safe, the backend must serialize access internally.

  • count_tokens must return a non-zero count for non-empty text. The registry probes this once at registration and rejects backends that report zero — a zero count would make every span appear to fit any budget. At runtime, a zero count for non-empty text is not trusted: the chunker substitutes the character count and logs the substitution. (An implementation may still return 0 for the empty string.)

  • count_tokens must not panic; return a best-effort count for text the tokenizer can’t fully process.

  • Counting should be deterministic for a given input — the splitter may evaluate overlapping spans of the same text repeatedly.

Count the tokens in text according to this backend’s tokenizer.

Signature:

uintptr_t xberg_tokenizer_backend_count_tokens(XBERGAlefHandle this, const char* text);

Example:

uintptr_t result = xberg_tokenizer_backend_count_tokens(instance, "value");

Parameters:

Name Type Required Description
text const char* Yes The text

Returns: uintptr_t


C representation: XBERGTranscriptionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTranscriptionConfig does not appear anywhere in the generated header.

Configuration for audio/video transcription (speech-to-text).

When present and enabled, Xberg will route audio and video files (mp3, mp4, m4a, wav, webm, etc.) through the transcription pipeline.

The heavy dependencies (ORT, hf-hub, symphonia) are only pulled when the transcription feature is enabled. The config struct itself is available under transcription-types so that ExtractionConfig round-trips on all targets.

All fields have sensible defaults. The recommended starting point is:

[extraction.transcription]
enabled = true
model = "tiny"
Field Type Default Description
enabled int32_t true Master switch. When false, the transcription pipeline is not run. The extractor is registered for audio/video MIME types whenever the transcription feature is compiled in, independently of this flag, so an audio/video input with enabled = false fails with an XbergError.Transcription explaining how to turn transcription on — it does not fall through to another extractor.
model XBERGAlefHandle XBERG_TINY Whisper model size to use. Smaller = faster + lower memory. tiny is the pragmatic default for first-time users and CI.
language const char* NULL Optional language hint (ISO-639-1 code, e.g. “en”, “de”). When NULL (default), the current engine falls back to English. For deterministic production output, always set this explicitly.
timestamps int32_t false Whether to request segment-level timestamps. When true, the decoder prompt omits <|notimestamps|> so the model emits <|x.xx|> tokens, and each transcript segment becomes its own paragraph element carrying start_ms / end_ms attributes. When false (default), all segment text is joined into a single flat paragraph with no timing attributes.
max_duration_ms uint64_t* 1800000 Hard safety limit on input duration (milliseconds). Files longer than this are rejected after decode, before model work. Default: 30 minutes. Set to NULL to disable (not recommended for untrusted input).
max_bytes uint64_t* 536870912 Hard safety limit on input size (bytes). Default: 512 MiB. Protects against pathological or malicious uploads.
timeout_ms uint64_t* 600000 Wall-clock timeout for the entire transcription operation (ms). Bounds audio decode, model resolution/download, and inference together. On expiry the extraction fails with an XbergError.Transcription. NULL disables the bound and lets the operation run unbounded (not recommended for untrusted input). Enforced on the async extraction path only; the size and duration caps (max_bytes, max_duration_ms) are checked on every path. Default: 10 minutes.
model_cache_dir const char* NULL Optional alternate Hugging Face cache root for Whisper models. When unset, hf-hub follows HF_HUB_CACHE, HUGGINGFACE_HUB_CACHE, HF_HOME, XDG, and platform defaults. Files remain in the standard content-addressed snapshot layout and are not copied into an Xberg cache.
allow_network int32_t true Allow network access to download models from Hugging Face Hub. When false, only previously cached models may be used. Useful for air-gapped or fully offline deployments.
verify_hash int32_t false Request SHA256 verification of downloaded model files. Defaults to false because the resolver downloads from mutable Hugging Face refs unless callers pin and verify models out-of-band. Explicit true requests are rejected by the model resolver until pinned checksum metadata is available.

C representation: XBERGTranslation is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTranslation does not appear anywhere in the generated header.

Translation of the extracted content.

Holds the translated rendition of ExtractedDocument.content and (when preserve_markup was requested) the translated formatted_content. Chunks are translated in place inside ExtractedDocument.chunks[*].content rather than duplicated here.

Field Type Default Description
target_lang const char* — BCP-47 language tag the translation was produced into (e.g. "de", "fr-CA").
source_lang const char* NULL BCP-47 source language. NULL when the translation backend was asked to detect.
content const char* — Translated plain-text body. Matches the shape of ExtractedDocument.content.
formatted_content const char* NULL Translated markup body (Markdown / HTML / etc.) when preserve_markup was enabled on the config. NULL otherwise.

Since: v1.0

C representation: XBERGTranslationConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTranslationConfig does not appear anywhere in the generated header.

Configuration for the translation post-processor.

Field Type Default Description
target_lang const char* — BCP-47 language tag for the target language (e.g. "de", "fr-CA").
source_lang const char* NULL Optional explicit source language. NULL asks the backend to auto-detect.
preserve_markup int32_t /* serde(default) */ Translate the formatted (Markdown/HTML) rendition alongside plain text when formatted_content is present.
llm XBERGAlefHandle — LLM configuration used for translation.

C representation: XBERGTreeSitterConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTreeSitterConfig does not appear anywhere in the generated header.

Configuration for tree-sitter language pack integration.

Controls grammar download behavior and code analysis options.

[tree_sitter]
languages = ["python", "rust"]
groups = ["web"]
[tree_sitter.process]
structure = true
comments = true
docstrings = true
Field Type Default Description
enabled int32_t true Enable code intelligence processing (default: true). When false, tree-sitter analysis is completely skipped even if the config section is present.
cache_dir const char* NULL Custom cache directory for downloaded grammars. When NULL, uses the default: ~/.cache/tree-sitter-language-pack/v{version}/libs/. Consumed both by the CLI (tree-sitter download --from-config, cache warm) and by CodeExtractor at extraction time, so that a configured cache directory is honoured wherever grammars are looked up or downloaded, not only during an explicit CLI download.
languages const char* NULL Languages to pre-download on init (e.g., ["python", "rust"]). Consumed only by the CLI’s tree-sitter download --from-config and cache warm commands as a pre-download hint. Extraction itself does not read this field: a given source file always processes with a single, already auto-detected language, so there is nothing for a language allowlist to gate at extraction time.
groups const char* NULL Language groups to pre-download (e.g., ["web", "systems", "scripting"]). Consumed only by the CLI’s tree-sitter download --from-config and cache warm commands, for the same reason as languages above.
process XBERGAlefHandle — Processing options for code analysis.

C representation: XBERGTreeSitterProcessConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGTreeSitterProcessConfig does not appear anywhere in the generated header.

Processing options for tree-sitter code analysis.

Controls which analysis features are enabled when extracting code files.

Field Type Default Description
structure int32_t true Extract structural items (functions, classes, structs, etc.). Default: true.
imports int32_t true Extract import statements. Default: true.
exports int32_t true Extract export statements. Default: true.
comments int32_t false Extract comments. Default: false.
docstrings int32_t false Extract docstrings. Default: false.
symbols int32_t false Extract symbol definitions. Default: false.
diagnostics int32_t false Include parse diagnostics. Default: false.
data_extraction int32_t false Extract a hierarchical key/value data tree from data-format files (JSON, YAML, TOML, XML, CSV, etc.). Default: false.
chunk_max_size uintptr_t* NULL Maximum chunk size in bytes. NULL disables chunking.
content_mode XBERGAlefHandle XBERG_CHUNKS Content rendering mode for code extraction.

C representation: XBERGUrlExtractionConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGUrlExtractionConfig does not appear anywhere in the generated header.

URL ingestion and crawl configuration.

Field Type Default Description
mode XBERGAlefHandle XBERG_AUTO URL extraction mode.
crawl XBERGAlefHandle xberg::UrlExtractionConfig::default_xberg_crawl_config() Crawlberg crawl configuration used for HTTP(S) URL extraction.
document_url_pattern const char* NULL Optional regex filter for document-discovered URLs.
max_document_urls_per_result uint32_t* 100 Maximum URLs to follow per extraction result.
max_total_urls uint32_t* 1000 Maximum URLs followed across the whole extraction call.
allow_local_file_inputs int32_t true Allow bare local filesystem path inputs.
allow_file_uris int32_t true Allow local file:// URI inputs.

C representation: XBERGUserChunkConfig is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGUserChunkConfig does not appear anywhere in the generated header.

User-provided chunk configuration.

Field Type Default Description
page_ranges const char* NULL User-specified page ranges (overrides automatic chunking).
pages_per_chunk uint32_t* NULL User-specified pages per chunk (overrides automatic calculation).
force_chunking int32_t — Force chunking even for small documents.
disable_chunking int32_t — Disable chunking even for large documents.

C representation: XBERGValidator is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGValidator does not appear anywhere in the generated header.

Trait for validator plugins.

Validators check extraction results for quality, completeness, or correctness. Unlike post-processors, validator errors fail fast - if a validator returns an error, the extraction fails immediately.

  • Quality Gates: Ensure extracted content meets minimum quality standards
  • Compliance: Verify content meets regulatory requirements
  • Content Filtering: Reject documents containing unwanted content
  • Format Validation: Verify extracted content structure
  • Security Checks: Scan for malicious content

Validator errors are fatal - they cause the extraction to fail and bubble up to the caller. Use validators for hard requirements that must be met.

For non-fatal checks, use post-processors instead.

Validators must be thread-safe (Send + Sync).

Validate an extraction result.

Check the extraction result and return Ok(()) if valid, or an error if validation fails.

Returns:

  • Ok(()) if validation passes
  • Err(...) if validation fails (extraction will fail)

Errors:

  • XbergError.Validation - Validation failed
  • Any other error type appropriate for the failure

Signature:

int32_t xberg_validator_validate(XBERGAlefHandle this, XBERGAlefHandle result, XBERGAlefHandle config);

Example:

xberg_validator_validate(instance, 0, 0);

Parameters:

Name Type Required Description
result XBERGAlefHandle Yes The extraction result to validate
config XBERGAlefHandle Yes Extraction configuration

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.

Optional: Check if this validator should run for a given result.

Allows conditional validation based on MIME type, metadata, or content. Defaults to true (always run).

Returns:

true if the validator should run, false to skip.

Signature:

int32_t xberg_validator_should_validate(XBERGAlefHandle this, XBERGAlefHandle result, XBERGAlefHandle config);

Example:

int32_t result = xberg_validator_should_validate(instance, 0, 0);

Parameters:

Name Type Required Description
result XBERGAlefHandle Yes The extracted document
config XBERGAlefHandle Yes The extraction config

Returns: int32_t

Optional: Get the validation priority.

Higher priority validators run first. Useful for ordering validation checks (e.g., run cheap validations before expensive ones).

Default priority is 50.

Returns:

Priority value (higher = runs earlier).

Signature:

int32_t xberg_validator_priority(XBERGAlefHandle this);

Example:

int32_t result = xberg_validator_priority(instance);

Returns: int32_t


C representation: XBERGXlsxAppProperties is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGXlsxAppProperties does not appear anywhere in the generated header.

Application properties from docProps/app.xml for XLSX

Contains Excel-specific document metadata.

Field Type Default Description
application const char* NULL Application name (e.g., “Microsoft Excel”)
app_version const char* NULL Application version
doc_security int32_t* NULL Document security level
scale_crop int32_t* NULL Scale crop flag
links_up_to_date int32_t* NULL Links up to date flag
shared_doc int32_t* NULL Shared document flag
hyperlinks_changed int32_t* NULL Hyperlinks changed flag
company const char* NULL Company name
worksheet_names const char* NULL Worksheet names

C representation: XBERGXmlExtractionResult is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGXmlExtractionResult does not appear anywhere in the generated header.

XML extraction result.

Contains extracted text content from XML files along with structural statistics about the XML document.

Field Type Default Description
content const char* — Extracted text content (XML structure filtered out)
element_count uintptr_t — Total number of XML elements processed
unique_elements const char* — List of unique element names found (sorted)

C representation: XBERGXmlMetadata is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGXmlMetadata does not appear anywhere in the generated header.

XML metadata extracted during XML parsing.

Provides statistics about XML document structure.

Field Type Default Description
element_count uint32_t — Total number of XML elements processed
unique_elements const char* NULL List of unique element tag names (sorted)

C representation: XBERGYakeParams is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGYakeParams does not appear anywhere in the generated header.

YAKE-specific parameters.

Field Type Default Description
window_size uintptr_t 2 Window size for co-occurrence analysis (default: 2). Controls the context window for computing co-occurrence statistics.

C representation: XBERGYearRange is a documentation-only name for this type. The C ABI hands you a scalar XBERGAlefHandle handle – the literal string XBERGYearRange does not appear anywhere in the generated header.

Year range for bibliographic metadata.

Field Type Default Description
min uint32_t* NULL Earliest (minimum) year in the range.
max uint32_t* NULL Latest (maximum) year in the range.
years const char* /* serde(default) */ All individual years present in the collection.

ONNX Runtime execution provider type.

Determines which hardware backend is used for model inference. Auto (default) selects the best available provider per platform.

Value Description
XBERG_AUTO Auto-select: CoreML on macOS, CUDA on Linux, CPU elsewhere.
XBERG_CPU CPU execution provider (always available).
XBERG_CORE_ML Apple CoreML (macOS/iOS Neural Engine + GPU).
XBERG_CUDA NVIDIA CUDA GPU acceleration.
XBERG_TENSOR_RT NVIDIA TensorRT (optimized CUDA inference).

Since: v1.1

Selects which evidence is authoritative when Xberg infers a MIME type.

Value Description
XBERG_PREFER_CONTENT Prefer a supported content signature, falling back to the filename extension.
XBERG_TRUST_EXTENSION Trust a supported filename extension without reading content for detection. # Security Filenames are attacker-controlled input in uploads and downloads. Use this only for trusted sources; otherwise a misleading extension can route arbitrary content to the wrong extractor.
XBERG_CONTENT_ONLY Ignore the filename extension and require content-based detection.

Target format for re-encoding extracted images.

Controls whether and how extracted images are normalised to a uniform container format before being returned in ExtractedDocument.images. The default (Native) preserves the format produced by each extractor without any additional encode pass.

Callers that need uniform output — e.g. cloud pipelines that always store WebP thumbnails — set this once on ImageExtractionConfig.output_format rather than re-encoding downstream.

Uses a tagged enum: {"type": "native"}, {"type": "png"}, {"type": "jpeg", "quality": 90}, etc.

Value Description
XBERG_NATIVE Preserve whatever format the extractor produced (default). No re-encode pass is performed. ExtractedImage.format reflects the source format: JPEG for embedded PDF images, PNG for rasterised content, or the native container format from office documents.
XBERG_PNG Re-encode all extracted images as PNG (lossless).
XBERG_JPEG Re-encode all extracted images as JPEG at the given quality level. quality must be in 1..=100. Values outside this range are clamped and a warning is emitted. Higher values produce larger files with less artefacting; 85 is a reasonable default. — Fields: quality: uint8_t
XBERG_WEBP Re-encode all extracted images as WebP at the given quality level. quality must be in 1..=100. Values outside this range are clamped and a warning is emitted. 80 is a reasonable default. — Fields: quality: uint8_t
XBERG_HEIF Re-encode all extracted images as HEIF/HEIC at the given quality level. quality must be in 1..=100. Values outside this range are clamped and a warning is emitted. 80 is a reasonable default. The encode path requires the heic feature; on builds without it, selecting this variant returns an EncodeFailed warning and leaves the image untouched. — Fields: quality: uint8_t
XBERG_SVG Output pure-vector SVG. Lossless. Raster sources are not re-encoded (a warning is emitted and the image bytes are left untouched). When the source is already SVG, the bytes are passed through the usvg sanitizer (strips external hrefs, JS event handlers, and foreignObject elements) when SvgOptions.sanitize is true. Requires the svg feature.

Source kind for ExtractInput.

Value Description
XBERG_BYTES Raw in-memory bytes.
XBERG_URI A filesystem path, file:// URI, or HTTP(S) URL.

URL extraction mode.

Value Description
XBERG_AUTO Classify HTTP(S) resources after fetch.
XBERG_DOCUMENT Treat the URI as a single remote document/page.
XBERG_CRAWL Crawl from the seed URI and extract discovered pages/documents.

Output format for extraction results.

Controls the format of the content field in ExtractedDocument. When set to Markdown, Djot, or Html, the output uses that format. Plain returns the raw extracted text.

Value Description
XBERG_PLAIN Plain text content only (default)
XBERG_MARKDOWN Markdown format
XBERG_DJOT Djot markup format
XBERG_HTML HTML format
XBERG_JSON JSON tree format with heading-driven sections.
XBERG_DOC_TAGS Docling DocTags format (tables rendered as OTSL).
XBERG_CUSTOM Custom renderer registered via the RendererRegistry. The string is the renderer name (e.g., “docx”, “latex”). — Fields: 0: const char*

Controls how Jupyter notebook code cells are rendered during extraction.

A code cell carries both its source and any outputs that were saved in the notebook. Callers ingesting notebooks for AI agents want different slices of this depending on the task. Xberg never executes cells — Outputs and Both only surface outputs already stored in the .ipynb.

This toggle governs a code cell’s source body and its saved outputs. Markdown (prose) cells and structural markers (kernel language, cell id, tags, execution count) are unaffected — prose always renders and markers orient the reader regardless of mode.

Value Description
XBERG_SOURCE Render the code source as a fenced code block; omit saved outputs.
XBERG_OUTPUTS Omit the code source; render only the saved cell outputs.
XBERG_BOTH Render both the code source and the saved outputs (default; preserves the historical behavior).

Built-in HTML theme selection.

Value Description
XBERG_DEFAULT Sensible defaults: system font stack, neutral colours, readable line measure. CSS custom properties (--kb-*) are all defined so user CSS can override individual values.
XBERG_GIT_HUB GitHub Markdown-inspired palette and spacing.
XBERG_DARK Dark background, light text.
XBERG_LIGHT Minimal light theme with generous whitespace.
XBERG_UNSTYLED No built-in stylesheet emitted. CSS custom properties are still defined on :root so user stylesheets can reference var(--kb-*) tokens.

Late-interaction model types supported by Xberg.

Value Description
XBERG_PRESET Use a preset ColBERT model (recommended). — Fields: name: const char*
XBERG_CUSTOM Use a custom ColBERT ONNX model from HuggingFace. — Fields: model_id: const char*, model_file: const char*, additional_files: const char*, max_length: int64_t
XBERG_PLUGIN In-process late-interaction backend registered via the plugin system. — Fields: name: const char*

Formula recognition model selection.

Value Description
XBERG_LATEX_OCR RapidLaTeXOCR (MIT, pix2tex-derived): resizer + encoder + decoder ONNX, ~180 MB total, downloaded on demand.

Which table structure recognition model to use.

Controls the model used for table cell detection within layout-detected table regions. Wire format is snake_case in all serializers (JSON, TOML, YAML).

Value Description
XBERG_TATR TATR (Table Transformer) – default, 30MB, DETR-based row/column detection.
XBERG_SLANET_WIRED SLANeXT wired variant – 365MB, optimized for bordered tables.
XBERG_SLANET_WIRELESS SLANeXT wireless variant – 365MB, optimized for borderless tables.
XBERG_SLANET_PLUS SLANet-plus – 7.78MB, lightweight general-purpose.
XBERG_SLANET_AUTO Classifier-routed SLANeXT: auto-select wired/wireless per table. Uses PP-LCNet classifier (6.78MB) + both SLANeXT variants (730MB total).
XBERG_DISABLED Disable table structure model inference entirely; use heuristic path only.

How to resolve overlapping native vs layout (TATR/SLANeXT) tables.

When both native detection and the layout table model produce a table for the same page region, one must be dropped. This controls which one wins. Wire format is snake_case in all serializers (JSON, TOML, YAML).

Value Description
XBERG_CONTENT Keep whichever table carries more content (cell count + markdown length). This is the historical default. TATR/SLANeXT tables usually recognize more cells and therefore win, which maximizes table-structure F1 but can lower text F1 when the recognized cell reflow diverges from the source reading order.
XBERG_NATIVE Prefer the native table when it overlaps a layout table. Native tables preserve the source reading order, which scores higher on text F1 for documents where the layout model’s cell reflow diverges from the ground truth.
XBERG_LAYOUT Prefer the layout (TATR/SLANeXT) table when it overlaps a native table.

Which PDF pages the layout model runs on.

Layout detection renders each selected page to a raster and runs ONNX inference on it, which dominates extraction cost. This controls page selection; LayoutStrategy.Always preserves the historical behavior of running on every page. Wire format is snake_case in all serializers (JSON, TOML, YAML).

Value Description
XBERG_ALWAYS Run layout detection unconditionally on every page.
XBERG_AUTO Pre-screen each page with cheap geometry signals and run the model only on pages likely to benefit (multi-column, table-bearing, figure-heavy, form-like, or rotated pages). Pages the pre-screen skips are processed exactly like pages where the model ran and found no regions. On the OCR path only inference is skipped; page rasters are still produced because OCR consumes them. For non-PDF inputs Auto behaves as LayoutStrategy.Always.

Since: v1.1

Managed credential-provider configuration for OAuth2/STS-based authentication modes liter-llm cannot express via a static api_key. See LlmConfig.credential_provider.

Debug is implemented by hand: CredentialProviderConfig.AzureAd’s client_secret is a credential and must never be printed, matching LlmConfig’s own redaction policy. The other variants carry no secret material — CredentialProviderConfig.VertexOauth2 and CredentialProviderConfig.BedrockWebIdentity reference a file path, never the key or token itself.

Value Description
XBERG_AZURE_AD Azure AD OAuth2 client-credentials flow (Azure OpenAI / Azure Cognitive Services). — Fields: tenant_id: const char*, client_id: const char*, client_secret: const char*, scope: const char*
XBERG_VERTEX_OAUTH2 Google Vertex AI OAuth2 via a service-account JSON key file on disk. Points at a file path rather than embedding the key inline: the key file contains an RSA private key — stronger secret material than an API key — and LlmConfig must never carry that directly, matching the credential-handling policy the rest of this module follows. — Fields: service_account_key_file: const char*, scope: const char*
XBERG_VERTEX_ADC Google Vertex AI Application Default Credentials, resolved from the GCE/GKE/Cloud Run metadata server. Carries no secret material at all. — Fields: scope: const char*
XBERG_BEDROCK_WEB_IDENTITY AWS STS AssumeRoleWithWebIdentity (EKS IRSA / OIDC federation) for Bedrock. — Fields: role_arn: const char*, token_file: const char*, session_name: const char*, region: const char*

How a structured-extraction preset is dispatched to the model.

This is the preset-facing call mode (the preferred_call_mode field of a Preset). The structured pipeline has a richer runtime-only decision enum with skip and fallback states; this 3-variant type is the stable, serializable surface presets and bindings depend on.

Value Description
XBERG_TEXT_ONLY Use the extracted text only.
XBERG_VISION_ONLY Use rasterized page images only.
XBERG_TEXT_PLUS_VISION Provide both extracted text and page images to the model.

How partial results from multiple model calls (e.g. per page batch) are combined.

Canonical home for the merge strategy referenced by presets and by the structured pipeline’s post-processing. There is intentionally only one merge type across the crate — do not introduce a second.

Value Description
XBERG_OBJECT_MERGE Deep-merge JSON objects field by field (later calls fill missing fields).
XBERG_ARRAY_CONCAT Concatenate top-level arrays across calls.
XBERG_OBJECT_FIRST Keep the first non-empty result; ignore subsequent calls.

NER backend selector.

Value Description
XBERG_ONNX xberg-gliner ONNX inference. Requires ner-onnx feature. Models download lazily from xberg-io/gliner-models.
XBERG_LLM liter-llm zero-shot NER via structured-output prompts. Requires ner-llm feature. Useful when domain-specific categories outstrip the ONNX taxonomy.

Policy controlling when VLM (Vision Language Model) OCR is used as a fallback.

This knob is syntactic sugar over the explicit OcrPipelineConfig stage ordering. When vlm_fallback is set and pipeline is NULL, an equivalent pipeline is synthesised at extraction time:

  • VlmFallbackPolicy.Disabled — no synthesis; single-backend mode (default).

  • VlmFallbackPolicy.OnLowQuality — tries the classical backend first; if the result scores below quality_threshold, tries VLM.

  • VlmFallbackPolicy.Always — skips the classical backend and sends every page to the VLM.

When OcrConfig.pipeline is explicitly set, vlm_fallback is ignored — the explicit pipeline takes precedence.

Errors:

Both OnLowQuality and Always require OcrConfig.vlm_config to be Some. Constructing an OcrConfig with one of these policies but no vlm_config is detected by OcrConfig.validate and will surface as a Validation error at extraction time, not a panic.

Value Description
XBERG_DISABLED No VLM fallback (default). Behaves identically to the pre-policy single-backend mode.
XBERG_ON_LOW_QUALITY Try the classical OCR backend first. If the quality score is below quality_threshold, send the page to the VLM. quality_threshold is in the [0.0, 1.0] range, but it is not the same quantity reported on score (GH#1584). The accept decision blends text-shape quality with confidence, weighted 0.7/0.3 (extractors.pdf.ocr.pipeline_stage_score) – when the backend’s confidence is on a known scale it contributes only 30% of the compared score, so a page can clear this threshold on clean-looking text even while its own PageOcrConfidence.score reads below it. Do not calibrate this value by reading PageOcrConfidence.score off a sample page and expecting an equal quality_threshold to reproduce the same accept/reject outcome. A value of 0.5 is a reasonable starting point; calibrate with the Stage 0 benchmark harness. — Fields: quality_threshold: double
XBERG_ALWAYS Skip the classical OCR backend entirely. Every page is sent to the VLM.

Which pages of a PDF get OCR’d when neither force_ocr nor force_ocr_pages applies.

Value Description
XBERG_AUTO OCR only when the native text layer fails a quality check (default). A scanner’s invisible OCR sidecar passes that check, so scanned pages carrying one are extracted natively. Use OcrStrategy.ScannedPages to OCR them instead.
XBERG_SCANNED_PAGES Additionally OCR every page that looks like a scan. Pages are graded on raster coverage, whether the text layer is invisible or absent, the image codec, and the producer. Pages at or above min_confidence are OCR’d; the rest keep native text and still go through the Auto quality check. Detects that a text layer came from a scanner, not whether it is accurate, so a page carrying a good sidecar is OCR’d too. — Fields: min_confidence: double

PDF extraction backend selection.

Controls which engine parses and renders PDF documents. Wire format is snake_case in all serializers (JSON, TOML, YAML). Defaults to PdfBackend.Native – selecting anything else never changes behavior for a caller who does not opt in.

Value Description
XBERG_NATIVE xberg’s own pure-Rust PDF engine (default), crates/xberg-native-pdf.
XBERG_PDFIUM pdfium – Google’s PDFium engine, gated behind the pdf-pdfium Cargo feature (#700 added the selection level; #702 added the extraction engine, extractors.pdf.pdfium_engine, deliberately smaller in scope than Native – see that module’s doc comment for exactly what it extracts). A build without the pdf-pdfium feature rejects this at the CLI validation layer rather than silently falling back to Native. A build with the feature still enforces the selection at extraction time (extractors.pdf.PdfExtractor.extract_core), because that is the one dispatch point every caller – CLI, library use, API/MCP servers, language bindings – passes through; see that function’s doc comment for why the enforcement lives there and not only in CLI validation.

Controls how markdown tables are handled when they exceed the chunk size limit.

Only applies when chunker_type is Markdown.

  • Split - Default behavior: tables are split at row boundaries like any other block element. Continuation chunks contain only data rows without the header, which can break downstream consumers that need column context.

  • RepeatHeader - Prepend the table header (header row + separator row) to every continuation chunk that contains data rows from the same table. Adds a small amount of duplicate text but ensures each chunk is self-contained for extraction, search, and LLM consumption.

Value Description
XBERG_SPLIT Split tables at row boundaries (default). Continuation chunks have no header.
XBERG_REPEAT_HEADER Prepend the table header to every chunk that continues a split table.

Type of text chunker to use.

  • Text - Generic text splitter, splits on whitespace and punctuation
  • Markdown - Markdown-aware splitter, preserves formatting and structure
  • Yaml - YAML-aware splitter, creates one chunk per top-level key
  • Semantic - Topic-aware chunker. With an EmbeddingConfig, splits at embedding-based topic shifts tuned by topic_threshold (default 0.75, lower = more splits). Without an embedding, falls back to a structural-boundary heuristic (ALL-CAPS headers, numbered sections, blank-line paragraphs) and merges groups into chunks capped at max_characters (default 1000). topic_threshold has no effect in the fallback path. For best results, pair with an embedding model.
Value Description
XBERG_TEXT Generic whitespace- and punctuation-aware text splitter (default).
XBERG_MARKDOWN Markdown-aware splitter that preserves heading and code-block boundaries.
XBERG_YAML YAML-aware splitter that creates one chunk per top-level key.
XBERG_SEMANTIC Topic-aware chunker that splits at embedding-based topic shifts.

How chunk size is measured.

Defaults to Characters (Unicode character count). When using token-based sizing, chunks are sized by token count according to the specified tokenizer.

Token-based sizing uses HuggingFace tokenizers loaded at runtime, or a tokenizer backend you register yourself. Any tokenizer available on HuggingFace Hub can be used, including OpenAI-compatible tokenizers (e.g., Xenova/gpt-4o, Xenova/cl100k_base). To size chunks with your own tokenizer instead (llama.cpp/GGUF vocabularies, SentencePiece models, custom vocabs), register a TokenizerBackend with register_tokenizer_backend and set model to the registered name.

Value Description
XBERG_CHARACTERS Size measured in Unicode characters (default).
XBERG_TOKENIZER Size measured in tokens from a HuggingFace tokenizer or a registered tokenizer backend. — Fields: model: const char*, cache_dir: const char*

Embedding model types supported by Xberg.

Value Description
XBERG_PRESET Use a preset model configuration (recommended) — Fields: name: const char*
XBERG_CUSTOM Use a custom ONNX model from HuggingFace — Fields: model_id: const char*, dimensions: uintptr_t
XBERG_LLM Provider-hosted embedding model via liter-llm. Uses the model specified in the nested LlmConfig (e.g., "openai/text-embedding-3-small"). — Fields: llm: XBERGAlefHandle
XBERG_PLUGIN In-process embedding backend registered via the plugin system. The caller registers an EmbeddingBackend once (e.g. a wrapper around an already-loaded llama-cpp-python, sentence-transformers, or tuned ONNX model), then references it by name in config. Xberg calls back into the registered backend during chunking and standalone embed requests — no HuggingFace download, no ONNX Runtime requirement, no HTTP sidecar. When this variant is selected, only the following EmbeddingConfig fields apply: normalize (post-call L2 normalization) and max_embed_duration_secs (dispatcher timeout). Model-loading fields (batch_size, cache_dir, show_download_progress, acceleration) are ignored — the host owns the model lifecycle, so there is no download to report progress for. Semantic chunking falls back to ChunkingConfig.max_characters when this variant is used, since there is no preset to look a chunk-size ceiling up against — size your context window via max_characters directly. See register_embedding_backend. — Fields: name: const char*

Selects how a local ONNX reranker’s raw output tensor is turned into a score.

  • RerankerHead.CrossEncoder — classic single-logit cross-encoder head: the model emits [batch, 1] (or [batch]) logits; the caller applies sigmoid to get a [0, 1] score. This is the original, unchanged path.

  • RerankerHead.Qwen3Generative — Qwen3 generative-reranker head: the model emits [batch, seq, vocab] logits; the score is P("yes") read from the last token’s logits over the “yes”/“no” vocabulary entries, via a softmax over those two logits. Already a [0, 1] probability — no sigmoid is applied.

Value Description
XBERG_CROSS_ENCODER Single-logit cross-encoder head (sigmoid applied by the caller).
XBERG_QWEN3_GENERATIVE Qwen3 generative-reranker head (softmax over yes/no token logits).

Reranker model types supported by Xberg.

Value Description
XBERG_PRESET Use a preset cross-encoder model (recommended). — Fields: name: const char*
XBERG_CUSTOM Use a custom ONNX cross-encoder from HuggingFace. — Fields: model_id: const char*, model_file: const char*, additional_files: const char*, max_length: int64_t, head: XBERGAlefHandle
XBERG_LLM Provider-hosted reranker via liter-llm (e.g. Cohere, Jina, Voyage). The model in the nested LlmConfig must be a rerank-capable model ID (e.g. "cohere/rerank-english-v3.0"). — Fields: llm: XBERGAlefHandle
XBERG_PLUGIN In-process reranker registered via the plugin system. The caller registers a RerankerBackend once (e.g. a wrapper around a sentence-transformers cross-encoder or a provider client), then references it by name in config. Xberg calls back into the registered backend — no HuggingFace download, no ONNX Runtime requirement. When this variant is selected, only max_rerank_duration_secs applies. Model-loading fields (batch_size, cache_dir, show_download_progress, acceleration) are ignored — the host owns the model lifecycle, so there is no download to report progress for. See register_reranker_backend. — Fields: name: const char*

Sparse-embedding model types supported by Xberg.

Value Description
XBERG_PRESET Use a preset SPLADE model (recommended). — Fields: name: const char*
XBERG_CUSTOM Use a custom SPLADE (BertForMaskedLM) ONNX model from HuggingFace. — Fields: model_id: const char*, model_file: const char*, additional_files: const char*, max_length: int64_t
XBERG_PLUGIN In-process sparse-embedding backend registered via the plugin system. — Fields: name: const char*

Supported Whisper model sizes.

These map to published ONNX exports on Hugging Face (onnx-community or similar orgs). The actual filenames and repos are resolved inside the transcription engine.

Value Description
XBERG_TINY Smallest, fastest, lowest quality. Good default for development and CI.
XBERG_BASE Reasonable quality/speed tradeoff.
XBERG_SMALL Better accuracy with higher memory and cache use.
XBERG_MEDIUM High quality; slower and more memory-intensive.
XBERG_LARGE_V3 Best quality (large-v3). Use only when latency and memory use are acceptable.

Content rendering mode for code extraction.

Controls how extracted code content is represented in the content field of ExtractedDocument.

Value Description
XBERG_CHUNKS Use TSLP semantic chunks as content (default).
XBERG_RAW Use raw source code as content.
XBERG_STRUCTURE Emit function/class headings + docstrings (no code bodies).

Type of list detection.

Value Description
XBERG_BULLET Bullet points (-, *, •, etc.)
XBERG_NUMBERED Numbered lists (1., 2., etc.)
XBERG_LETTERED Lettered lists (a., b., A., B., etc.)
XBERG_INDENTED Indented items

OCR backend types.

Value Description
XBERG_TESSERACT Tesseract OCR (native Rust binding)
XBERG_PADDLE_OCR PaddleOCR (Python-based, via FFI)
XBERG_CANDLE Candle-based VLM OCR (TrOCR, PaddleOCR-VL).
XBERG_CUSTOM Name-selected built-in or third-party OCR backend.

How a backend’s reported page-level confidence must be interpreted.

Backend confidence scores are not interchangeable. Tesseract’s mean word confidence is a classifier score validated to track legibility on a 0-100 scale. Sceptre (EasyOCR-based) reports a length-penalised custom_mean that is rescaled into the same 0-100 range but is not comparable — its ordering can be inverted relative to legibility (a dense prose page can score lower than a nearly-blank one). A page-rejection gate calibrated on Tesseract’s scale was once applied unconditionally to sceptre’s output and rejected every page of a 16-page document, emptying it. This descriptor exists so gating code can ask a backend what its number means instead of assuming.

Value Description
XBERG_LEGIBILITY Validated to track legibility on a known scale — usable as an absolute quality gate. — Fields: scale_max: double
XBERG_UNCALIBRATED A number is reported, but it is not validated to correlate with legibility. Never gate on it.
XBERG_NONE No page-level confidence is reported at all.

How a backend copes with a page raster whose text is not upright.

Rotated-page handling is a backend capability, not a universal guarantee. An A/B run this session against /Rotate 270 scanned pages showed the three handled cases genuinely differ: Tesseract reconstructs correct reading order on a sideways raster outright; PaddleOCR recognises the rotated text correctly (it warps each detected quad upright before running recognition) but leaves its block list in raw raster (y, x) order, so the caller must reorder; sceptre produces character garbage on the same sideways raster and only reads correctly once the page is rendered upright first. A caller that skips an upright-render step for a backend that actually needs one gets silent garbage, not an error.

Read this before “simplifying” the type. There is exactly one decision point in the codebase that inspects this value: upright_raster_for_backend (crate.extractors.pdf.ocr), which tests orientation_handling != RequiresUpright and otherwise does nothing. Every other mention forwards the value to that test. So, to that codepath, SelfCorrecting and RecognisesRotatedText are behaviourally identical — the enum is a boolean at the point of use, and the three variants describe measured backend behaviour rather than three dispatch paths.

RecognisesRotatedText’s actual remedy is not this enum. The block-order fix is the backend_options["page_rotation_degrees"] hint injected by ocr_config_with_page_rotation_hint (crate.extractors.pdf.ocr) unconditionally, for every backend, which PaddleOcrBackend.process_image reads back (page_rotation_degrees_from_backend_options -> residual_rotation_for_reorder -> reorder_blocks_for_page_rotation, crate.paddle_ocr.backend) and applies internally. Declaring RecognisesRotatedText therefore changes nothing on its own; a backend in that class must also read the hint. Conversely, gating that hint on this enum would remove a field from Tesseract’s OcrConfig and hence from the OCR cache key (hash(image + language + config)), invalidating every cached page — do not do it without its own A/B.

Only the PDF OCR routes call OcrBackend.page_orientation_handling (the --force-ocr route via extract_with_ocr and the scanned-pages route via extract_mixed_ocr_native). The raw-image route (crate.extractors.image) never calls it: there is no /Rotate to consult, and orientation there is handled by the PP-LCNet document-orientation classifier (crate.doc_orientation) gated on OcrConfig.auto_rotate.

The trait default is RequiresUpright (deliberately the least capable option, see OcrBackend.page_orientation_handling). A backend that does not declare therefore pays, on every page with /Rotate != 0, a re-encode plus rotation of the raster in upright_raster_for_backend and a bounding-box round-trip back through undo_upright_raster_correction. That is the safe direction to be wrong in, but it is not free, and for a backend that never got measured it is not known to be necessary either.

Value Description
XBERG_SELF_CORRECTING Reconstructs reading order regardless of page rotation — safe to hand a raster in any orientation.
XBERG_RECOGNISES_ROTATED_TEXT Recognises rotated text correctly but emits blocks in raw raster order, so the caller must reorder.
XBERG_REQUIRES_UPRIGHT Requires an upright raster; rotated text produces garbage.

Processing stages for post-processors.

Post-processors are executed in stage order (Early → Middle → Late). Use stages to control the order of post-processing operations.

Value Description
XBERG_EARLY Early stage - foundational processing. Use for: - Language detection - Character encoding normalization - Entity extraction (NER) - Text quality scoring
XBERG_MIDDLE Middle stage - content transformation. Use for: - Keyword extraction - Token reduction - Text summarization - Semantic analysis
XBERG_LATE Late stage - final enrichment. Use for: - Custom user hooks - Analytics/logging - Final validation - Output formatting

Intensity level for the token-reduction pipeline.

Value Description
XBERG_OFF No reduction applied; text is returned as-is.
XBERG_LIGHT Remove only the most common stopwords.
XBERG_MODERATE Balanced stopword removal and redundancy filtering.
XBERG_AGGRESSIVE Aggressive filtering; may remove less common content words.
XBERG_MAXIMUM Maximum compression; prioritizes brevity over completeness.

Type of PDF annotation.

Value Description
XBERG_TEXT Sticky note / text annotation
XBERG_HIGHLIGHT Highlighted text region
XBERG_LINK Hyperlink annotation
XBERG_STAMP Rubber stamp annotation
XBERG_UNDERLINE Underline text markup
XBERG_STRIKE_OUT Strikeout text markup
XBERG_SQUIGGLY Squiggly (wavy) underline text markup — Since: v1.1
XBERG_INK Freehand drawing (ink) annotation — Since: v1.1
XBERG_SQUARE Rectangle/box shape annotation — Since: v1.1
XBERG_CIRCLE Ellipse/oval shape annotation — Since: v1.1
XBERG_POLYGON Closed polygon shape annotation — Since: v1.1
XBERG_POLY_LINE Open polyline shape annotation — Since: v1.1
XBERG_LINE Line annotation — Since: v1.1
XBERG_CARET Caret (text-insertion marker) annotation — Since: v1.1
XBERG_FILE_ATTACHMENT Embedded file attachment annotation — Since: v1.1
XBERG_SOUND Embedded sound annotation — Since: v1.1
XBERG_MOVIE Embedded movie annotation — Since: v1.1
XBERG_OTHER Any other annotation type

Types of block-level elements in Djot.

Value Description
XBERG_PARAGRAPH Standard prose paragraph.
XBERG_HEADING Section heading (level stored in FormattedBlock.level).
XBERG_BLOCKQUOTE Block quotation container.
XBERG_CODE_BLOCK Fenced or indented code block.
XBERG_LIST_ITEM Individual item within a list.
XBERG_ORDERED_LIST Numbered (ordered) list container.
XBERG_BULLET_LIST Unnumbered (bullet) list container.
XBERG_TASK_LIST Task / checkbox list container.
XBERG_DEFINITION_LIST Definition list container.
XBERG_DEFINITION_TERM Term part of a definition list entry.
XBERG_DEFINITION_DESCRIPTION Description / definition part of a definition list entry.
XBERG_DIV Generic div container with optional attributes.
XBERG_SECTION Logical section container, often associated with a heading.
XBERG_THEMATIC_BREAK Horizontal rule / thematic break.
XBERG_RAW_BLOCK Raw content block in a specified format (e.g. HTML, LaTeX).
XBERG_MATH_DISPLAY Display-mode mathematical expression.

Types of inline elements in Djot.

Value Description
XBERG_TEXT Plain text run.
XBERG_STRONG Bold / strong emphasis.
XBERG_EMPHASIS Italic / regular emphasis.
XBERG_HIGHLIGHT Highlighted text (marker pen).
XBERG_SUBSCRIPT Subscript text.
XBERG_SUPERSCRIPT Superscript text.
XBERG_INSERT Inserted text (tracked change).
XBERG_DELETE Deleted text (tracked change).
XBERG_CODE Inline code span.
XBERG_LINK Hyperlink with URL.
XBERG_IMAGE Inline image reference.
XBERG_SPAN Generic inline span with optional attributes.
XBERG_MATH Inline mathematical expression.
XBERG_RAW_INLINE Raw inline content in a specified format.
XBERG_FOOTNOTE_REF Footnote reference marker.
XBERG_SYMBOL Named symbol or emoji shortcode.

Semantic kind of a relationship between document elements.

Value Description
XBERG_FOOTNOTE_REFERENCE Footnote marker -> footnote definition.
XBERG_CITATION_REFERENCE Citation marker -> bibliography entry.
XBERG_INTERNAL_LINK Internal anchor link (#id) -> target heading/element.
XBERG_CAPTION Caption paragraph -> figure/table it describes.
XBERG_LABEL Label -> labeled element (HTML <label for>, LaTeX \label{}).
XBERG_TOC_ENTRY TOC entry -> target section.
XBERG_CROSS_REFERENCE Cross-reference (LaTeX \ref{}, DOCX cross-reference field).

Content layer classification for document nodes.

Replaces separate body/furniture arrays with per-node granularity.

Value Description
XBERG_BODY Main document body content.
XBERG_HEADER Page/section header (running header).
XBERG_FOOTER Page/section footer (running footer).
XBERG_FOOTNOTE Footnote content.

Tagged enum for node content. Each variant carries only type-specific data.

Uses #[serde(tag = "node_type")] to avoid “type” keyword collision in Go/Java/TypeScript bindings.

Value Description
XBERG_TITLE Document title. — Fields: text: const char*
XBERG_HEADING Section heading with level (1-6). — Fields: level: uint8_t, text: const char*
XBERG_PARAGRAPH Body text paragraph. — Fields: text: const char*
XBERG_LIST List container — children are ListItem nodes. — Fields: ordered: int32_t
XBERG_LIST_ITEM Individual list item. — Fields: text: const char*
XBERG_TABLE Table with structured cell grid. — Fields: grid: XBERGAlefHandle
XBERG_IMAGE Image reference. — Fields: description: const char*, image_index: uint32_t, src: const char*
XBERG_CODE Code block. — Fields: text: const char*, language: const char*
XBERG_QUOTE Block quote — container, children carry the quoted content.
XBERG_FORMULA Mathematical formula / equation. — Fields: text: const char*
XBERG_FOOTNOTE Footnote reference content. — Fields: text: const char*
XBERG_COMMENT Reviewer/editor comment content (e.g. DOCX comments). Distinct from NodeContent.Footnote (xberg-io/xberg#300): comments and footnotes both reach the internal document via a marker/definition pair, but a consumer needs to tell a reviewer comment apart from an authored footnote. — Fields: text: const char* — Since: v1.1
XBERG_GROUP Logical grouping container (section, key-value area). heading_level + heading_text capture the section heading directly rather than relying on a first-child positional convention. — Fields: label: const char*, heading_level: uint8_t, heading_text: const char*
XBERG_PAGE_BREAK Page break marker.
XBERG_SLIDE Presentation slide container — children are the slide’s content nodes. — Fields: number: uint32_t, title: const char*
XBERG_DEFINITION_LIST Definition list container — children are DefinitionItem nodes.
XBERG_DEFINITION_ITEM Individual definition list entry with term and definition. — Fields: term: const char*, definition: const char*
XBERG_CITATION Citation or bibliographic reference. — Fields: key: const char*, text: const char*
XBERG_ADMONITION Admonition / callout container (note, warning, tip, etc.). Children carry the admonition body content. — Fields: kind: const char*, title: const char*
XBERG_RAW_BLOCK Raw block preserved verbatim from the source format. Used for content that cannot be mapped to a semantic node type (e.g. JSX in MDX, raw LaTeX in markdown, embedded HTML). — Fields: format: const char*, content: const char*
XBERG_METADATA_BLOCK Structured metadata block (email headers, YAML frontmatter, etc.). — Fields: entries: const char*

Types of inline text annotations.

Value Description
XBERG_BOLD Bold (strong) text formatting.
XBERG_ITALIC Italic (emphasis) text formatting.
XBERG_UNDERLINE Underlined text.
XBERG_STRIKETHROUGH Strikethrough text.
XBERG_CODE Inline code span.
XBERG_SUBSCRIPT Subscript text.
XBERG_SUPERSCRIPT Superscript text.
XBERG_LINK Hyperlink annotation. — Fields: url: const char*, title: const char*
XBERG_HIGHLIGHT Highlighted text (PDF highlights, HTML <mark>).
XBERG_COLOR Text color (CSS-compatible value, e.g. “#ff0000”, “red”). — Fields: value: const char*
XBERG_FONT_SIZE Font size with units (e.g. “12pt”, “1.2em”, “16px”). — Fields: value: const char*
XBERG_CUSTOM Extensible annotation for format-specific styling. — Fields: name: const char*, value: const char*

Standard entity categories produced by built-in NER backends.

The Custom(String) variant lets caller-supplied categories (e.g. LLM schemas) flow through without losing fidelity to the consumer.

Value Description
XBERG_PERSON A person’s name.
XBERG_ORGANIZATION A company, institution, or organisation name.
XBERG_LOCATION A geographic location (city, country, address).
XBERG_DATE A calendar date.
XBERG_TIME A time of day or duration.
XBERG_MONEY A monetary amount with optional currency.
XBERG_PERCENT A percentage value.
XBERG_EMAIL An email address.
XBERG_PHONE A phone number.
XBERG_URL A URL or URI.
XBERG_CUSTOM A caller-supplied custom category label. — Fields: 0: const char*

How the extracted text was produced.

Value Description
XBERG_NATIVE Text extracted directly from the document’s native format (no OCR).
XBERG_OCR All text was obtained via OCR (e.g. scanned image-only PDF).
XBERG_MIXED Text came from a combination of native extraction and OCR.

Semantic structural classification of a text chunk.

Assigned by the heuristic classifier in chunking.classifier. Defaults to Unknown when no rule matches. Designed to be extended in future versions without breaking changes.

Value Description
XBERG_HEADING Section heading or document title.
XBERG_PARTY_LIST Party list: names, addresses, and signatories.
XBERG_DEFINITIONS Definition clause (“X means…”, “X shall mean…”).
XBERG_OPERATIVE_CLAUSE Operative clause containing legal/contractual action verbs.
XBERG_SIGNATURE_BLOCK Signature block with signatures, names, and dates.
XBERG_SCHEDULE Schedule, annex, appendix, or exhibit section.
XBERG_TABLE_LIKE Table-like content with aligned columns or repeated patterns.
XBERG_FORMULA Mathematical formula or equation.
XBERG_CODE_BLOCK Code block or preformatted content.
XBERG_FUNCTION Function or method definition (tree-sitter structured code chunking).
XBERG_CLASS Class, struct, interface, or trait definition (tree-sitter structured code chunking).
XBERG_MODULE Module, namespace, or top-level file scope (tree-sitter structured code chunking).
XBERG_IMAGE Embedded or referenced image content.
XBERG_ORG_CHART Organizational chart or hierarchy diagram.
XBERG_DIAGRAM Diagram, figure, or visual illustration.
XBERG_UNKNOWN Unclassified or mixed content.

Heuristic classification of what an image likely depicts.

Value Description
XBERG_PHOTOGRAPH Photographic image (natural scene, photograph)
XBERG_DIAGRAM Technical or schematic diagram
XBERG_CHART Chart, graph, or plot
XBERG_DRAWING Freehand or technical drawing
XBERG_TEXT_BLOCK Text-heavy image (scanned text, document)
XBERG_DECORATION Decorative element or border
XBERG_LOGO Logo or brand mark
XBERG_ICON Small icon
XBERG_TILE_FRAGMENT Fragment of a larger tiled image (tile of a technical drawing)
XBERG_MASK Mask or transparency map
XBERG_PAGE_RASTER Full-page render produced during OCR preprocessing; used as a citation thumbnail.
XBERG_UNKNOWN Could not classify with reasonable confidence

Result-shape selection for extraction results.

Distinct from OutputFormat (which controls rendering — Plain, Markdown, HTML, etc.). ResultFormat controls the shape of the result: a unified content blob vs. an element-based decomposition.

Value Description
XBERG_UNIFIED Unified format with all content in content field
XBERG_ELEMENT_BASED Element-based format with semantic element extraction

Semantic element type classification.

Categorizes text content into semantic units for downstream processing. Supports the element types commonly found in Unstructured documents.

Value Description
XBERG_TITLE Document title
XBERG_NARRATIVE_TEXT Main narrative text body
XBERG_HEADING Section heading
XBERG_LIST_ITEM List item (bullet, numbered, etc.)
XBERG_TABLE Table element
XBERG_IMAGE Image element
XBERG_PAGE_BREAK Page break marker
XBERG_CODE_BLOCK Code block
XBERG_FORMULA Mathematical formula (LaTeX source in text)
XBERG_BLOCK_QUOTE Block quote
XBERG_FOOTER Footer text
XBERG_HEADER Header text

Kind of a PDF form field.

Mirrors xberg_native_pdf’s widget field taxonomy without leaking the upstream type across the binding surface.

Value Description
XBERG_TEXT Single- or multi-line text input.
XBERG_CHECKBOX Checkbox (on/off toggle).
XBERG_RADIO Radio-button group member.
XBERG_CHOICE Choice field (dropdown or list box).
XBERG_SIGNATURE Digital-signature field.
XBERG_BUTTON Push button.
XBERG_UNKNOWN Field type that could not be classified.

Format-specific metadata (discriminated union).

Only one format type can exist per extraction result. This provides type-safe, clean metadata without nested optionals.

Value Description
XBERG_PDF Metadata extracted from a PDF document. — Fields (flattened into the tagged object): pdf_version: const char*, producer: const char*, is_encrypted: int32_t, width: int64_t, height: int64_t, page_count: uint32_t, scanned_confidence: float, scanned_pages: const char*, fabricated_text_pages: const char*, implausible_text_pages: const char*, layout_gated_pages: const char*, layout_gate_reasons: const char*
XBERG_DOCX Metadata extracted from a DOCX Word document. — Fields (flattened into the tagged object): core_properties: XBERGAlefHandle, app_properties: XBERGAlefHandle, custom_properties: const char*
XBERG_EXCEL Metadata extracted from an Excel spreadsheet. — Fields (flattened into the tagged object): sheet_count: uint32_t, sheet_names: const char*
XBERG_EMAIL Metadata extracted from an email message (EML/MSG). — Fields (flattened into the tagged object): from_email: const char*, from_name: const char*, to_emails: const char*, cc_emails: const char*, bcc_emails: const char*, message_id: const char*, attachments: const char*
XBERG_PPTX Metadata extracted from a PowerPoint presentation. — Fields (flattened into the tagged object): slide_count: uint32_t, slide_names: const char*, image_count: uint32_t, table_count: uint32_t
XBERG_ARCHIVE Metadata extracted from an archive (ZIP, TAR, 7Z, etc.). — Fields (flattened into the tagged object): format: const char*, file_count: uint32_t, file_list: const char*, total_size: uint64_t, compressed_size: uint64_t
XBERG_IMAGE Metadata extracted from a raster or vector image. — Fields (flattened into the tagged object): width: uint32_t, height: uint32_t, format: const char*, exif: const char*
XBERG_XML Metadata extracted from an XML document. — Fields (flattened into the tagged object): element_count: uint32_t, unique_elements: const char*
XBERG_TEXT Metadata extracted from a plain-text file. — Fields (flattened into the tagged object): line_count: uint32_t, word_count: uint32_t, character_count: uint32_t, headers: const char*, links: const char*, code_blocks: const char*
XBERG_HTML Metadata extracted from an HTML document. — Fields (flattened into the tagged object): title: const char*, description: const char*, keywords: const char*, author: const char*, canonical_url: const char*, base_href: const char*, language: const char*, text_direction: XBERGAlefHandle, open_graph: const char*, twitter_card: const char*, meta_tags: const char*, headers: const char*, links: const char*, images: const char*, structured_data: const char*
XBERG_OCR Metadata produced by an OCR pipeline. — Fields (flattened into the tagged object): language: const char*, psm: int32_t, output_format: const char*, table_count: uint32_t, table_rows: uint32_t, table_cols: uint32_t
XBERG_CSV Metadata extracted from a CSV or TSV file. — Fields (flattened into the tagged object): row_count: uint32_t, column_count: uint32_t, delimiter: const char*, has_header: int32_t, column_types: const char*
XBERG_BIBTEX Metadata extracted from a BibTeX bibliography file. — Fields (flattened into the tagged object): entry_count: uintptr_t, citation_keys: const char*, authors: const char*, year_range: XBERGAlefHandle, entry_types: const char*
XBERG_CITATION Metadata extracted from a citation file (RIS, PubMed, EndNote). — Fields (flattened into the tagged object): citation_count: uintptr_t, format: const char*, authors: const char*, year_range: XBERGAlefHandle, dois: const char*, keywords: const char*
XBERG_FICTION_BOOK Metadata extracted from a FictionBook (FB2) e-book. — Fields (flattened into the tagged object): genres: const char*, sequences: const char*, annotation: const char*
XBERG_DBF Metadata extracted from a dBASE (DBF) database file. — Fields (flattened into the tagged object): record_count: uintptr_t, field_count: uintptr_t, fields: const char*
XBERG_JATS Metadata extracted from a JATS (Journal Article Tag Suite) XML file. — Fields (flattened into the tagged object): copyright: const char*, license: const char*, history_dates: const char*, contributor_roles: const char*
XBERG_EPUB Metadata extracted from an EPUB e-book. — Fields (flattened into the tagged object): coverage: const char*, dc_format: const char*, relation: const char*, source: const char*, dc_type: const char*, cover_image: const char*
XBERG_PST Metadata extracted from an Outlook PST archive. — Fields (flattened into the tagged object): message_count: uintptr_t
XBERG_AUDIO Metadata extracted from an audio or video file. — Fields (flattened into the tagged object): duration_ms: uint64_t, codec: const char*, container: const char*, sample_rate_hz: uint32_t, channels: uint16_t, bitrate: uint32_t
XBERG_CODE Code (tree-sitter analyzable source). Carries the structural chunks (function, class, and module boundaries) produced by the tree-sitter extractor, consumed by the chunking pipeline to emit structure-aware Chunks instead of falling back to text-based splitting. Wraps CodeMetadata (a named struct) rather than const CodeChunkInfo* directly: FormatMetadata is internally tagged (#[serde(tag = "format_type")]), and serde cannot serialize a tagged newtype variant that wraps a sequence — the tag has no map to live in. Wrapping a struct gives serde a map to hold the tag, and keeps this variant shape consistent with every sibling (Variant(XMetadata)) so the derived OpenAPI discriminator can reference a named component schema. — Fields (flattened into the tagged object): chunks: const char*, data: XBERGAlefHandle

Discriminates the shape of a CodeDataNode.

Purpose-built mirror of tree_sitter_language_pack.DataNodeKind — kept as an xberg-owned type so binding generators never need to resolve the upstream crate’s types across FFI/language boundaries.

Value Description
XBERG_KEY_VALUE A key/value pair or mapping (JSON/TOML/properties/YAML/HCL/CUE/KDL pair, or a wrapper “object”/“mapping” container).
XBERG_ELEMENT An XML element with a tag name in key and attributes in attributes.
XBERG_SEQUENCE A positional sequence item (JSON array element, YAML block sequence item, CSV/PSV row or cell).

Text direction enumeration for HTML documents.

Value Description
XBERG_LEFT_TO_RIGHT Left-to-right text direction
XBERG_RIGHT_TO_LEFT Right-to-left text direction
XBERG_AUTO Automatic text direction detection

Link type classification.

Value Description
XBERG_ANCHOR Anchor link (#section)
XBERG_INTERNAL Internal link (same domain)
XBERG_EXTERNAL External link (different domain)
XBERG_EMAIL Email link (mailto:)
XBERG_PHONE Phone link (tel:)
XBERG_OTHER Other link type

Image type classification.

Value Description
XBERG_DATA_URI Data URI image
XBERG_INLINE_SVG Inline SVG
XBERG_EXTERNAL External image URL
XBERG_RELATIVE Relative path image

Structured data type classification.

Value Description
XBERG_JSON_LD JSON-LD structured data
XBERG_MICRODATA Microdata
XBERG_RDFA RDFa

Bounding geometry for an OCR element.

Supports both axis-aligned rectangles (from Tesseract) and 4-point quadrilaterals (from PaddleOCR and rotated text detection).

Value Description
XBERG_RECTANGLE Axis-aligned bounding box (typical for Tesseract output). — Fields: left: uint32_t, top: uint32_t, width: uint32_t, height: uint32_t
XBERG_QUADRILATERAL 4-point quadrilateral for rotated/skewed text (PaddleOCR). Points are in clockwise order starting from top-left: [top_left, top_right, bottom_right, bottom_left] — Fields: points: const char*

Hierarchical level of an OCR element.

Maps to Tesseract’s page segmentation hierarchy and provides equivalent semantics for PaddleOCR.

Value Description
XBERG_WORD Individual word
XBERG_LINE Line of text (default for PaddleOCR)
XBERG_BLOCK Paragraph or text block
XBERG_PAGE Page-level element

Type of paginated unit in a document.

Distinguishes between different types of “pages” (PDF pages, presentation slides, spreadsheet sheets).

Value Description
XBERG_PAGE Standard document pages (PDF, DOCX, images)
XBERG_SLIDE Presentation slides (PPTX, ODP)
XBERG_SHEET Spreadsheet sheets (XLSX, ODS)

Strategy applied when a PII match is rewritten.

Value Description
XBERG_MASK Replace the matched span with a fixed mask token (default "[REDACTED]").
XBERG_HASH Replace with a SHA-256 hash of the original value (truncated to 16 hex chars). Lets downstream consumers do equality joins without recovering the source.
XBERG_TOKEN_REPLACE Replace with a per-category running token ("[PERSON_1]", "[PERSON_2]", …) so the same person referenced twice gets the same token within the document.
XBERG_DROP Delete the matched span entirely.

PII categories the pattern engine recognises.

Value Description
XBERG_EMAIL Email address (e.g. user@example.com).
XBERG_PHONE Phone number in any common format.
XBERG_SSN US Social Security Number.
XBERG_CREDIT_CARD Payment card number (Visa, Mastercard, Amex, etc.).
XBERG_POSTAL_CODE Postal / ZIP code.
XBERG_IP_ADDRESS IPv4 or IPv6 address.
XBERG_IBAN International Bank Account Number.
XBERG_SWIFT_BIC SWIFT / BIC bank identifier code.
XBERG_DATE_OF_BIRTH Date of birth.
XBERG_PERSON Person name, surfaced by the optional NER backend.
XBERG_ORGANIZATION Organization name, surfaced by the optional NER backend.
XBERG_LOCATION Location, surfaced by the optional NER backend.
XBERG_CUSTOM Caller-supplied custom category (e.g. internal employee IDs). Surfaced by the redaction engine when a hit comes from RedactionConfig.custom_terms or RedactionConfig.custom_patterns. The string is the label passed alongside the term/pattern. Use those fields rather than constructing Custom directly via the categories filter — the pattern engine cannot detect arbitrary text from a category name alone. — Fields: 0: const char*

Classification of a detected layout region that warrants VLM extraction.

Value Description
XBERG_FIGURE A figure, diagram, chart, or image region.
XBERG_DENSE_TABLE A densely formatted or complex table.
XBERG_COMPLEX_LAYOUT A region with complex or mixed layout.
XBERG_CAPTION A standalone image to caption.

A single line in a unified-diff hunk.

Defined here (rather than only in crate.diff) so RevisionDelta can reference it unconditionally, without requiring the diff Cargo feature. crate.diff re-exports this type verbatim.

Value Description
XBERG_CONTEXT Unchanged context line. — Fields: text: const char*
XBERG_ADDED Line added in the “after” version. — Fields: text: const char*
XBERG_REMOVED Line removed from the “before” version. — Fields: text: const char*

Semantic classification of a tracked change.

Value Description
XBERG_INSERTION Text or content was inserted.
XBERG_DELETION Text or content was deleted.
XBERG_FORMAT_CHANGE Run-level formatting (font, size, colour, …) was changed.
XBERG_COMMENT A reviewer comment or annotation.

Best-effort document location for a revision.

Value Description
XBERG_PARAGRAPH Body paragraph, identified by its zero-based index in the document flow. — Fields: index: uintptr_t
XBERG_TABLE_CELL Cell inside a table. — Fields: row: uintptr_t, col: uintptr_t, table_index: uintptr_t
XBERG_PAGE Page, identified by its zero-based index. — Fields: index: uintptr_t
XBERG_SLIDE Presentation slide, identified by its zero-based index. — Fields: index: uintptr_t
XBERG_SHEET Spreadsheet cell or range, identified by sheet index and optional name. — Fields: index: uintptr_t, name: const char*

Summarisation strategy.

Value Description
XBERG_EXTRACTIVE Pure-Rust extractive summary (TextRank over the chunk graph). Deterministic, fast, no external service required.
XBERG_ABSTRACTIVE Abstractive summary produced by liter-llm. Requires liter-llm feature and a configured LlmConfig. Token usage is captured in ExtractedDocument.llm_usage.

Semantic classification of an extracted URI.

Value Description
XBERG_HYPERLINK A clickable hyperlink (web URL, file link).
XBERG_IMAGE An image or media resource reference.
XBERG_ANCHOR An internal anchor or cross-reference target.
XBERG_CITATION A citation or bibliographic reference (DOI, academic ref).
XBERG_REFERENCE A general reference (e.g. \ref{} in LaTeX, :ref: in RST).
XBERG_EMAIL An email address (mailto: link or bare email).

Inference backend that an EmbeddingPreset runs on.

Onnx presets require the embeddings feature (ONNX Runtime, not available on WASM/Android x86_64 emulator). Static presets require static-embeddings (pure-Rust model2vec inference, no ORT — the only dense-embedding backend available on no-ort-target).

Defaults to Onnx via #[serde(default)] so every existing preset payload (which predates this field) keeps deserializing without change.

Value Description
XBERG_ONNX ONNX Runtime transformer inference (the historical, default backend).
XBERG_STATIC Pure-Rust static (model2vec) inference — no ONNX Runtime.

Keyword algorithm selection.

Value Description
XBERG_YAKE YAKE (Yet Another Keyword Extractor) - statistical approach
XBERG_RAKE RAKE (Rapid Automatic Keyword Extraction) - co-occurrence based

Schema-validation outcome surfaced as one of three buckets.

Fold into the combined confidence score without leaking internal validation error types.

Value Description
XBERG_ALL_VALID Every batch validated against the schema.
XBERG_PARTIAL_VALID At least one batch validated; at least one did not.
XBERG_ALL_INVALID No batch validated.

Reason for not chunking a document.

Value Description
XBERG_SMALL_FILE File is below size threshold. — Fields: size_bytes: uint64_t, threshold_bytes: uint64_t
XBERG_FEW_PAGES Document has fewer pages than threshold. — Fields: page_count: uint32_t, threshold: uint32_t
XBERG_TEXT_LAYER_DETECTED PDF has substantial text layer (OCR not needed). — Fields: text_coverage: float, avg_chars_per_page: uint32_t
XBERG_FORMAT_NOT_CHUNKABLE Document format does not support chunking. — Fields: mime_type: const char*
XBERG_CHUNKING_DISABLED Chunking is disabled by configuration.
XBERG_FAST_TEXT_EXTRACTION Force OCR is disabled and text extraction is fast.

Reason for chunking a document.

Value Description
XBERG_LARGE_FILE File exceeds size threshold. — Fields: size_bytes: uint64_t, threshold_bytes: uint64_t
XBERG_MANY_PAGES Document has many pages. — Fields: page_count: uint32_t, threshold: uint32_t
XBERG_OCR_REQUIRED PDF requires OCR and is large. — Fields: page_count: uint32_t, force_ocr: int32_t
XBERG_LARGE_AND_MANY_PAGES Both size and page count exceed thresholds. — Fields: size_bytes: uint64_t, page_count: uint32_t

Reason for boundary detection.

Value Description
XBERG_START Start of PDF.
XBERG_PAGE_ONE_MARKER Page-one marker (“Page 1”, “1 of N”) detected.
XBERG_LETTERHEAD_RESET Letterhead reset after signature block.
XBERG_DENSITY_SHIFT Text density shift with low bigram overlap.
XBERG_END End of PDF.

High-level category used to group presets in the registry UI.

Value Description
XBERG_FINANCE Invoices, receipts, statements, purchase orders, W-9.
XBERG_IDENTITY Passports, drivers licenses, insurance cards.
XBERG_LEGAL Contracts, NDAs, agreements.
XBERG_LOGISTICS Bills of lading, customs declarations, packing lists.
XBERG_MEDICAL Clinical records, lab reports.
XBERG_HR Pay stubs, resumes, employment offers.
XBERG_OTHER Catch-all for documents that don’t fit the other categories.

Page Segmentation Mode for Tesseract OCR.

Value Description
XBERG_OSD_ONLY Orientation and script detection only.
XBERG_AUTO_OSD Automatic page segmentation with OSD.
XBERG_AUTO_ONLY Automatic page segmentation without OSD or OCR.
XBERG_AUTO Fully automatic page segmentation with no OSD (default).
XBERG_SINGLE_COLUMN Assume a single column of text of variable sizes.
XBERG_SINGLE_BLOCK_VERTICAL Assume a single uniform block of vertically aligned text.
XBERG_SINGLE_BLOCK Assume a single uniform block of text.
XBERG_SINGLE_LINE Treat the image as a single text line.
XBERG_SINGLE_WORD Treat the image as a single word.
XBERG_CIRCLE_WORD Treat the image as a single word in a circle.
XBERG_SINGLE_CHAR Treat the image as a single character.

Outcome of a single doctor check.

Value Description
XBERG_PASS The backend or setting will work as configured.
XBERG_WARN The check ran and found something actionable, but nothing is broken (e.g. stray cache files, stale model revisions). Never fails the report.
XBERG_FAIL The configured setup will not work (or will silently degrade) on this host.
XBERG_SKIP The check cannot run locally (e.g. model not cached, feature not compiled in); first real use decides, possibly after a download.

Which concrete ONNX inference engine PaddleOCR model loading uses.

Mirrors sceptre.Backend for the PaddleOCR backend: Ort is the native, full-featured path (acceleration/execution-provider hook, ONNX-embedded dictionary metadata); Tract is the pure-Rust, CPU-only path used on targets where ort cannot link (Android x86_64 emulator, WASM once wired).

Value Description
XBERG_ORT Native ONNX Runtime (requires the paddle-ocr-ort feature).
XBERG_TRACT Pure-Rust ONNX via tract (requires the paddle-ocr-tract feature).

Supported languages in PaddleOCR.

Maps user-friendly language codes to paddle-ocr-rs language identifiers.

Value Description
XBERG_ENGLISH English
XBERG_CHINESE Simplified Chinese
XBERG_JAPANESE Japanese
XBERG_KOREAN Korean
XBERG_GERMAN German
XBERG_FRENCH French
XBERG_LATIN Latin script (covers most European languages)
XBERG_CYRILLIC Cyrillic (Russian and related)
XBERG_TRADITIONAL_CHINESE Traditional Chinese
XBERG_THAI Thai
XBERG_GREEK Greek
XBERG_EAST_SLAVIC East Slavic (Russian, Ukrainian, Belarusian)
XBERG_ARABIC Arabic (Arabic, Persian, Urdu)
XBERG_DEVANAGARI Devanagari (Hindi, Marathi, Sanskrit, Nepali)
XBERG_TAMIL Tamil
XBERG_TELUGU Telugu

The 18 canonical document layout classes.

All model backends (RT-DETR, YOLO, etc.) map their native class IDs to this shared set. Models with fewer classes (DocLayNet: 11, PubLayNet: 5) map to the closest equivalent.

Wire format is snake_case in all serializers (JSON, TOML, YAML).

Value Description
XBERG_CAPTION Figure or table caption text.
XBERG_CHART Chart or graph visualization.
XBERG_FOOTNOTE Footnote or endnote text.
XBERG_FORMULA Mathematical formula or equation.
XBERG_LIST_ITEM A single item in a bulleted or numbered list.
XBERG_PAGE_FOOTER Running footer at the bottom of a page.
XBERG_PAGE_HEADER Running header at the top of a page.
XBERG_PICTURE Image, chart, or other graphical element.
XBERG_SECTION_HEADER Section heading.
XBERG_TABLE Data table.
XBERG_TEXT Body text paragraph.
XBERG_TITLE Document or chapter title.
XBERG_DOCUMENT_INDEX Table of contents or index.
XBERG_CODE Source code block.
XBERG_CHECKBOX_SELECTED Checkbox in selected state.
XBERG_CHECKBOX_UNSELECTED Checkbox in unselected state.
XBERG_FORM Form field or form element.
XBERG_KEY_VALUE_REGION Key-value pair region (e.g. label + value in a form).

Authentication configuration.

Value Description
XBERG_BASIC HTTP Basic authentication. — Fields: username: const char*, password: const char*
XBERG_BEARER Bearer token authentication. — Fields: token: const char*
XBERG_HEADER Custom authentication header. — Fields: name: const char*, value: const char*

When to use the headless browser fallback.

Value Description
XBERG_AUTO Automatically detect when JS rendering is needed and fall back to browser.
XBERG_ALWAYS Always use the browser for every request.
XBERG_NEVER Never use the browser fallback.
XBERG_STEALTH Always use the browser with all stealth surfaces enabled. Behaves like Always for escalation purposes (every request is routed through the browser tier), but additionally enables: - browser JavaScript stealth patches - native-backend TLS fingerprint spoofing - stealth-aware default user-agent when no explicit UA is set - 1920×1080 viewport override Use this instead of setting the now-removed BrowserConfig.stealth boolean field.

Wait strategy for browser page rendering.

Value Description
XBERG_NETWORK_IDLE Wait until network activity is idle.
XBERG_SELECTOR Wait for a specific CSS selector to appear in the DOM.
XBERG_FIXED Wait for a fixed duration after navigation.

Browser backend used for JavaScript rendering.

Value Description
XBERG_CHROMIUMOXIDE Existing Chromium/CDP backend powered by chromiumoxide.
XBERG_NATIVE Crawlberg-owned native browser backend derived from Obscura.

Opt-in encoding applied to a downloaded document’s bytes for callers who need the content available in a serializable field rather than reading it from disk.

NULL (the CrawlConfig.document_content_encoding default) produces neither — unlike screenshots, base64-encoding a document by default would duplicate an already up-to-document_max_size buffer (50 MB default) in memory per document.

Value Description
XBERG_BASE64 Populate DownloadedDocument.content_base64 with a base64-encoded copy.

Traversal order for a crawl.

Selects both the queue discipline and the selection strategy, because global order is a property of the frontier: the engine hands its bounded selection window to the strategy, so a strategy alone can only reorder URLs that have already been dequeued.

Value Description
XBERG_BFS Breadth-first: a FIFO frontier visits every URL at one depth before the next.
XBERG_DFS Depth-first: a LIFO frontier descends into a page’s children before its siblings.
XBERG_BEST_FIRST Highest-priority-first within the selection window, scored by CrawlStrategy.score_url.
XBERG_ADAPTIVE Like BestFirst, but stops once newly crawled pages stop contributing new terms.

Content filter applied to each crawled page before it reaches the result.

Value Description
XBERG_BM25 Keep only pages scoring at or above bm25_threshold for bm25_query.

The category of a downloaded asset.

Value Description
XBERG_DOCUMENT A document file (PDF, DOC, etc.).
XBERG_IMAGE An image file.
XBERG_AUDIO An audio file.
XBERG_VIDEO A video file.
XBERG_FONT A font file.
XBERG_STYLESHEET A CSS stylesheet.
XBERG_SCRIPT A JavaScript file.
XBERG_ARCHIVE An archive file (ZIP, TAR, etc.).
XBERG_DATA A data file (JSON, XML, CSV, etc.).
XBERG_OTHER An unrecognized asset type.

Hostname/IP allowlist matcher for SSRF policy.

Serializes as an internally-tagged object so each variant is distinguishable on the wire and round-trips losslessly:

{"type": "exact", "value": "api.example.com"}
{"type": "suffix", "value": ".example.com"}
{"type": "cidr", "value": "10.0.0.0/8"}

A bare JSON string is still accepted on deserialization and resolves to Exact, preserving configs written against the previous untagged representation.

Exact: HostMatcher.Exact

Value Description
XBERG_EXACT Exact hostname match (case-insensitive). — Fields: value: const char*
XBERG_SUFFIX Suffix match: “.xberg.io” matches “api.xberg.io” and “xberg.io”. — Fields: value: const char*
XBERG_CIDR CIDR match: “10.0.0.0/8” matches IP addresses in that range. — Fields: value: const char*

Controls which conversion tier is used.

Value Description
XBERG_AUTO Automatically pick the best tier for the input (default). Runs the classifier against the prescan report and uses Tier-1 when eligible; falls back to Tier-2 on bail or when the classifier routes to Tier-2.
XBERG_TIER2 Always use the Tier-2 (tl.parse + walk) path, skipping Tier-1.

HTML preprocessing aggressiveness level.

Controls the extent of cleanup performed before conversion. Higher levels remove more elements.

Value Description
XBERG_MINIMAL Minimal cleanup. Remove only essential noise (scripts, styles).
XBERG_STANDARD Standard cleanup. Default. Removes navigation, forms, and other auxiliary content.
XBERG_AGGRESSIVE Aggressive cleanup. Remove extensive non-content elements and structure.

Heading style options for Markdown output.

Controls how headings (h1-h6) are rendered in the output Markdown.

Value Description
XBERG_UNDERLINED Underlined style (=== for h1, — for h2).
XBERG_ATX ATX style (# for h1, ## for h2, etc.). Default.
XBERG_ATX_CLOSED ATX closed style (# title #, with closing hashes).

List indentation character type.

Controls whether list items are indented with spaces or tabs.

Value Description
XBERG_SPACES Use spaces for indentation. Default. Width controlled by list_indent_width.
XBERG_TABS Use tabs for indentation.

Whitespace handling strategy during conversion.

Determines how sequences of whitespace characters (spaces, tabs, newlines) are processed.

Value Description
XBERG_NORMALIZED Collapse multiple whitespace characters to single spaces. Default. Matches browser behavior.
XBERG_STRICT Preserve all whitespace exactly as it appears in the HTML.

Line break syntax in Markdown output.

Controls how soft line breaks (from <br> or line breaks in source) are rendered.

Value Description
XBERG_SPACES Two trailing spaces at end of line. Default. Standard Markdown syntax.
XBERG_BACKSLASH Backslash at end of line. Alternative Markdown syntax.

Code block fence style in Markdown output.

Determines how code blocks (<pre><code>) are rendered in Markdown.

Value Description
XBERG_INDENTED Indented code blocks (4 spaces). CommonMark standard.
XBERG_BACKTICKS Fenced code blocks with triple backticks. Default (GFM). Supports language hints.
XBERG_TILDES Fenced code blocks with tildes (~~~). Supports language hints.

Highlight rendering style for <mark> elements.

Controls how highlighted text is rendered in Markdown output.

Value Description
XBERG_DOUBLE_EQUAL Double equals syntax (==text==). Default. Pandoc-compatible.
XBERG_HTML Preserve as HTML (==text==). Original HTML tag.
XBERG_BOLD Render as bold (text). Uses strong emphasis.
XBERG_NONE Strip formatting, render as plain text. No markup.

Link rendering style in Markdown output.

Controls whether links and images use inline text syntax or reference-style [text][1] syntax with definitions collected at the end.

Value Description
XBERG_INLINE Inline links: text. Default.
XBERG_REFERENCE Reference-style links: [text][1] with [1]: url at end of document.

URL encoding strategy for link and image destinations.

Controls how special characters in URL destinations are handled when they require escaping to produce valid Markdown.

The Angle variant (default) wraps the destination in angle brackets: [text](<url with spaces>). This is the CommonMark-specified escape hatch but breaks when the URL itself contains >.

The Percent variant percent-encodes every character that is not an RFC 3986 unreserved character or /, producing a destination safe for all Markdown parsers: [text](url%20with%20spaces).

Value Description
XBERG_ANGLE Wrap destinations that contain spaces or newlines in angle brackets. Default.
XBERG_PERCENT Percent-encode all characters that are not RFC 3986 unreserved or /.

Output format for conversion.

Specifies the target markup language format for the conversion output.

Value Description
XBERG_MARKDOWN Standard Markdown (CommonMark compatible). Default.
XBERG_DJOT Djot lightweight markup language.
XBERG_PLAIN Plain text output (no markup, visible text only).

Main error type for all Xberg operations.

All errors in Xberg use this enum, which preserves error chains and provides context for debugging.

  • Io - File system and I/O errors (always bubble up)
  • Parsing - Document parsing errors (corrupt files, unsupported features)
  • Ocr - OCR processing errors
  • Validation - Input validation errors (invalid paths, config, parameters)
  • Cache - Cache operation errors (non-fatal, can be ignored)
  • ImageProcessing - Image manipulation errors
  • Serialization - JSON/MessagePack serialization errors
  • MissingDependency - Missing optional dependencies (tesseract, etc.)
  • Plugin - Plugin-specific errors
  • LockPoisoned - Mutex/RwLock poisoning (should not happen in normal operation)
  • UnsupportedFormat - Unsupported MIME type or file format
  • Other - Catch-all for uncommon errors
FFI error codes — a stable public contract
Section titled “FFI error codes — a stable public contract”

Each variant carries #[cfg_attr(alef, alef(error_code = N))]. Without it alef has no stable numeric taxonomy to bind to, so alef_ffi_error_code() collapses every variant to a single unknown code and EVERY language binding loses the ability to tell error kinds apart – errors.Is(err, ErrOcr) in Go, the checkLastError switch in Java, and Zig’s typed error sets all degrade to one catch-all. Refusing to guess is correct on alef’s part; supplying the codes is our job.

The blocks are XbergError 1000-1017, HeuristicsError 1100-1101, LoadError 1200-1205, ResolveError 1300, chosen to leave room to grow and to stay clear of alef’s own reserved 0-4 (NULL/Conversion/Unknown/Panic/InvalidHandle).

These numbers cross the FFI boundary and are compiled into released bindings, so they are append-only: never renumber or reuse a code, and give a new variant the next free number in its block rather than inserting one in declaration order.

Variant Description
XBERG_IO A file system or I/O operation failed. These errors always bubble up unchanged.
XBERG_PARSING Document parsing failed (e.g. corrupt file, unsupported format feature).
XBERG_OCR An OCR engine returned an error or produced unusable output.
XBERG_VALIDATION Invalid configuration or input parameters were supplied.
XBERG_CACHE A cache read or write operation failed.
XBERG_IMAGE_PROCESSING An image manipulation operation (resize, decode, DPI conversion) failed.
XBERG_SERIALIZATION JSON or MessagePack serialization/deserialization failed.
XBERG_MISSING_DEPENDENCY A required optional system dependency (e.g. tesseract) was not found.
XBERG_PLUGIN A registered plugin returned an error during extraction.
XBERG_LOCK_POISONED An internal Mutex or RwLock was found in a poisoned state.
XBERG_UNSUPPORTED_FORMAT The document’s MIME type is not supported by any registered extractor.
XBERG_EMBEDDING The embedding model or embedding pipeline returned an error.
XBERG_RERANKING The reranker model or reranking pipeline returned an error.
XBERG_TRANSCRIPTION Audio/video transcription failed.
XBERG_TIMEOUT The extraction operation exceeded the configured time limit.
XBERG_CANCELLED The extraction stopped after a cooperative cancellation request. Rust callers can request cancellation through CancellationToken. The REST async-jobs API uses the same mechanism. Generated language bindings do not currently expose in-process cancellation.
XBERG_SECURITY A security policy was violated (e.g. zip bomb, oversized archive).
XBERG_OTHER A catch-all for uncommon errors that do not fit another variant.

Errors that can occur during heuristics analysis.

Variant Description
XBERG_CONFIG_ERROR Invalid configuration value.
XBERG_PDF_ANALYSIS_ERROR PDF analysis step failed (only when heuristics-pdf feature is active).

Errors produced while loading or validating a preset file.

Variant Description
XBERG_PARSE The file is not valid JSON.
XBERG_SCHEMA_VALIDATION The file parses as JSON but does not validate against the meta-schema.
XBERG_DESERIALIZE The file validates but cannot be deserialized into Preset.
XBERG_ID_MISMATCH The preset’s declared id does not match its file-system location.
XBERG_BAD_META_SCHEMA The meta-schema itself failed to compile.
XBERG_IO A filesystem I/O error occurred while reading a preset directory.

Errors produced while resolving a preset against caller overrides.

Variant Description
XBERG_SCHEMA_NOT_OBJECT A custom schema override was supplied but is not a JSON object.

Trait Register Unregister Clear
OcrBackend xberg_register_ocr_backend xberg_unregister_ocr_backend xberg_clear_ocr_backend
PostProcessor xberg_register_post_processor xberg_unregister_post_processor xberg_clear_post_processor
Validator xberg_register_validator xberg_unregister_validator xberg_clear_validator
DocumentExtractor xberg_register_document_extractor xberg_unregister_document_extractor xberg_clear_document_extractor
EmbeddingBackend xberg_register_embedding_backend xberg_unregister_embedding_backend xberg_clear_embedding_backend
Renderer xberg_register_renderer xberg_unregister_renderer xberg_clear_renderer
RerankerBackend xberg_register_reranker_backend xberg_unregister_reranker_backend xberg_clear_reranker_backend
TokenizerBackend xberg_register_tokenizer_backend xberg_unregister_tokenizer_backend xberg_clear_tokenizer_backend