Output Formats
Choose the format that matches your downstream processing. See the Configuration Reference for all output format options.
Available Formats
Section titled “Available Formats”- Content Format —
output_formatcontrols thecontentfield: Plain, Markdown, Djot, HTML, JSON, DocTags, or a registered custom renderer such as DOT - Unified (default) — Plain text/Markdown, for LLM prompts and full-text search
- Element-Based — Flat array of typed elements with metadata, for RAG chunking and semantic search
- Document Structure — Hierarchical tree with explicit parent-child references, for knowledge graphs and structured apps
- PDF Hierarchy — Font-size classification into heading levels (H1–H6) for PDFs
- Image Output Formats — Normalize extracted images to PNG, JPEG, WebP, HEIF, or SVG
Unified Output (Default)
Section titled “Unified Output (Default)”No configuration required. The result contains:
content— Full document text with minimal formattingpages— Per-page breakdown for PDFs, DOCX, and PPTXtables— Extracted tables in structured formatimages— Image metadata and paths
Table Identity and Anchors
Section titled “Table Identity and Anchors”Every Table in tables[] (and in pages[].tables) carries a table_id and columns, so consumers can reconcile a table’s markdown block in rendered content with its structured entry:
table_id(str | None) — a stable identifier the extraction pipeline assigns deterministically (for example"table-3", in document order), so the same input document always produces the same ids. Same-page fragments of one physical table are already merged into a singletables[]entry before ids are assigned, so in practicetable_idis unique per entry rather than shared across several. A table split across a page boundary is not linked today — its per-page pieces get separate ids; sharing one id across page-boundary fragments is a documented future extension.Nonewhen the extractor did not assign one.columns(list[str] | None) — the header cells (the first row ofcells), populated even when the fragment’s own header row was merged away or lives in a sibling fragment. This makes a single fragment interpretable in isolation without needingtable_idto look up siblings.Nonewhen no header row could be determined.
Set table_anchors=True on ExtractionConfig to have the renderer additionally emit a [TABLE:{table_id}] marker immediately before each table’s Markdown block in content, pages[].content, and chunks[].content — a lightweight anchor a consumer can regex for to splice structured table data into rendered output. Only takes effect when output_format is Markdown or Djot. Defaults to False, so existing output is byte-identical unless explicitly enabled.
from xberg import ExtractInput, ExtractionConfig, extract
config = ExtractionConfig(output_format="markdown", table_anchors=True)output = await extract(ExtractInput(kind="uri", uri="report.pdf"), config=config)result = output.results[0]# result.content contains "[TABLE:table-1]" immediately before that table's markdown.for table in result.tables or []: print(table.table_id, table.columns)use xberg::{extract, ExtractInput, ExtractionConfig, OutputFormat};
let config = ExtractionConfig { output_format: OutputFormat::Markdown, table_anchors: true, ..Default::default()};let output = extract(ExtractInput::from_uri("report.pdf"), &config).await?;let result = &output.results[0];for table in &result.tables { println!("{:?} {:?}", table.table_id, table.columns);}output_format = "markdown"table_anchors = trueContent Output Format
Section titled “Content Output Format”output_format sets the format of the content field. It is independent of result_format (element-based / document structure). For OCR-backed scanned PDFs, non-Plain formats invoke the same geometry-derived structure heuristic (headings, lists) as native PDF extraction — but invoking it is not the same as it succeeding, and how much it recovers depends on the OCR backend. Measured on a 16-page scanned reference fixture with no layout detection, Tesseract promoted both headings and list items, Sceptre promoted few, and PaddleOCR promoted no headings at all. Enabling optional ML layout detection raises heading counts on every backend, but lowers list-item accuracy against ground truth. output_format != Plain is necessary but not always sufficient for structured OCR output.
Structure contract. output_format = Plain skips the structure heuristic entirely and returns raw
text. Any other output_format always runs the heuristic; when layout detection is also enabled, its
detections are layered on top. Layout detection is off by default, and Plain is the default
output_format. The two are not currently linked: enabling layout detection while output_format
stays Plain still runs the layout model, but its result is computed and then discarded rather than
forcing a non-Plain output.
OutputFormat |
Description |
|---|---|
Plain (default) |
Raw extracted text with minimal formatting |
Markdown |
Markdown-formatted content |
Djot |
Djot markup |
Html |
HTML-formatted content |
Json |
JSON tree with heading-driven sections |
DocTags |
Docling DocTags tag-stream format (tables as OTSL). See DocTags Output |
dot |
Graphviz DOT for a diagram recovered from a vector SVG or PDF source; empty string when no diagram was recovered. See Diagram DOT Output |
Custom(name) |
Output from a renderer registered in the RendererRegistry (e.g. docx, latex) |
config = ExtractionConfig(output_format="markdown")result = extract("document.pdf", config=config)print(result.content) # Markdown-formattedImage Output Formats
Section titled “Image Output Formats”Normalize extracted images to a uniform format after extraction but before post-processors.
By default, images are returned in their native format (JPEG from PDFs, PNG from rasterization, etc.). Set ImageExtractionConfig.output_format to re-encode all images to a single target format. This is useful for cloud pipelines that require uniform storage, thumbnails, or downstream processing.
Supported Formats
Section titled “Supported Formats”| Format | Quality param | Use case | Notes |
|---|---|---|---|
Native |
— | Default; preserve source format | No re-encode pass. Fastest. |
Png |
— | Lossless archival | Large file sizes; recommended for quality-critical workflows. |
Jpeg |
1–100 |
Web/cloud storage | Default quality 85. Lossy; good balance of size and quality. |
Webp |
1–100 |
Modern web use | Default quality 80. Better compression than JPEG; requires browser support. |
Heif |
1–100 |
Apple ecosystem | Default quality 80. Requires heic feature. Superior compression ratio vs JPEG/WebP. |
Svg |
— | Archival of vector images | Lossless vector output. Raster sources return a warning; not auto-vectorized. Requires svg feature. |
SVG Support
Section titled “SVG Support”When the svg feature is active and output_format is set to Svg:
- SVG → SVG: Re-parses the source via
usvgand re-serializes. Whensvg.sanitize = true(default), strips<script>elements, externalxlink:href/hrefattributes,<foreignObject>containers, and JavaScript event handlers. This is a lossy normalization for security. - SVG → PNG/JPEG/WebP/HEIF: Rasterizes to pixel format using
resvgat the specifiedrender_dpi(default 96.0, clamped 1.0–600.0 DPI). - Raster → SVG: Returns
EncodeWarning::UnsupportedDirection; bytes are left untouched. No auto-vectorization. - Security: SVG input capped at 10 MB; rasterized output capped at 16384² pixels (~1 GB peak). External resource loading is disabled.
Configuration
Section titled “Configuration”from xberg import ExtractionConfig, ImageExtractionConfig, ImageOutputFormat
# Normalize all images to WebP at quality 80config = ExtractionConfig( images=ImageExtractionConfig( output_format=ImageOutputFormat.Webp(quality=80) ))import { ExtractionConfig, ImageOutputFormat } from "@xberg-io/xberg";
const config: ExtractionConfig = { images: { outputFormat: { type: "webp", quality: 80 } }};use xberg::{ExtractionConfig, ImageExtractionConfig, ImageOutputFormat};
let config = ExtractionConfig { images: Some(ImageExtractionConfig { output_format: ImageOutputFormat::Webp { quality: 80 }, ..Default::default() }), ..Default::default()};SVG Sanitization
Section titled “SVG Sanitization”Enable or disable SVG security filtering:
from xberg import ExtractionConfig, ImageExtractionConfig, ImageOutputFormat, SvgOptions
# Re-encode SVG with sanitization disabledconfig = ExtractionConfig( images=ImageExtractionConfig( output_format=ImageOutputFormat.Svg, svg=SvgOptions(sanitize=False, render_dpi=96.0) ))use xberg::{ExtractionConfig, ImageExtractionConfig, ImageOutputFormat, SvgOptions};
let config = ExtractionConfig { images: Some(ImageExtractionConfig { output_format: ImageOutputFormat::Svg, svg: Some(SvgOptions { sanitize: false, render_dpi: 96.0 }), ..Default::default() }), ..Default::default()};Element-Based Output
Section titled “Element-Based Output”A flat array of typed elements (titles, paragraphs, tables, list items, code blocks, images, etc.). Each carries a page number; PDF text elements also carry bounding boxes when hierarchy extraction is enabled.
Use for RAG chunking, semantic search, or Unstructured.io-compatible pipelines.
Enable
Section titled “Enable”Tests element-based result format with element type assertions on DOCX
import asynciofrom xberg import extract, ExtractInput, ExtractInputKindfrom xberg._xberg import ExtractionConfig
async def main() -> None: input = ExtractInput(kind=ExtractInputKind("uri"), uri="https://example.com/docx/unit_test_headers.docx") config = ExtractionConfig.from_json("{\"result_format\":\"element_based\"}") result = await extract(input, config) for element in result.results[0].elements or []: print(element.element_type) print(element.text)
asyncio.run(main())Tests element-based result format with element type assertions on DOCX
import { ExtractInput, ExtractInputKind, ExtractionConfig, ResultFormat, extract } from "@xberg-io/xberg";async function main() { const input: ExtractInput = { kind: ExtractInputKind.Uri, uri: "https://example.com/docx/unit_test_headers.docx" }; const config: ExtractionConfig = { resultFormat: ResultFormat.ElementBased }; const result = await extract(input, config); const [first] = result.results ?? []; for (const element of first?.elements ?? []) { console.log(element.elementType); console.log(element.text); }}
void main();Tests element-based result format with element type assertions on DOCX
import { WasmExtractInput, WasmExtractInputKind, extract } from "@xberg-io/xberg-wasm";async function main() { const input: WasmExtractInput = (() => { const _u0 = WasmExtractInput.default(); _u0.kind = WasmExtractInputKind.Uri; _u0.uri = "https://example.com/docx/unit_test_headers.docx"; return _u0; })(); const result = await extract(input, { resultFormat: "element_based" }); const [first] = result.results ?? []; for (const element of first?.elements ?? []) { console.log(element.elementType); console.log(element.text); }}
void main();Tests element-based result format with element type assertions on DOCX
use xberg::extract;use xberg::ExtractInput;
#[tokio::main]async fn main() { let input_json: serde_json::Value = serde_json::from_str(r#"{"kind":"uri","uri":"https://example.com/docx/unit_test_headers.docx"}"#).unwrap(); let input = serde_json::from_value::<ExtractInput>(input_json).unwrap(); let config_json: serde_json::Value = serde_json::from_str(r#"{"result_format":"element_based"}"#).unwrap(); let config = serde_json::from_value(config_json).unwrap(); let result = extract(input, &config).await.expect("call failed"); for element in result.results[0].elements.iter().flatten() { println!("{:?}", element.element_type); println!("{:?}", element.text); }}Tests element-based result format with element type assertions on DOCX
package main
import ( "fmt" xberg "github.com/xberg-io/xberg/packages/go")
func ptr[T any](value T) *T { return &value }func main() { input := xberg.ExtractInput{ Kind: ptr(xberg.ExtractInputKindURI), URI: ptr(`https://example.com/docx/unit_test_headers.docx`), } config := xberg.ExtractionConfig{ ResultFormat: ptr(xberg.ResultFormatElementBased), } result, err := xberg.Extract(input, config) if err != nil { panic(err) } for _, element := range result.Results[0].Elements { fmt.Printf("%+v\n", element.ElementType) fmt.Printf("%+v\n", element.Text) }}Tests element-based result format with element type assertions on DOCX
import io.xberg.*;
public final class Example { public static void main(String[] args) throws Exception { var inputJson = "{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/unit_test_headers.docx\"}"; var input = JsonUtil.fromJson(inputJson, ExtractInput.class); var configJson = "{\"result_format\":\"element_based\"}"; var config = JsonUtil.fromJson(configJson, ExtractionConfig.class); var result = Xberg.extract(input, config); for (var element : result.results().get(0).elements()) { System.out.println(element.elementType()); System.out.println(element.text()); } }}Tests element-based result format with element type assertions on DOCX
import io.xberg.*import com.fasterxml.jackson.module.kotlin.jacksonObjectMapper
fun main() = kotlinx.coroutines.runBlocking { val mapper = jacksonObjectMapper().setPropertyNamingStrategy(com.fasterxml.jackson.databind.PropertyNamingStrategies.SNAKE_CASE) val input = mapper.readValue("{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/unit_test_headers.docx\"}", ExtractInput::class.java) val config = mapper.readValue("{\"result_format\":\"element_based\",\"url\":{\"crawl\":{\"ssrf\":{}}}}", ExtractionConfig::class.java) val result = Xberg.extract(input, config) for (element in result.results.first().elements.orEmpty()) { println(element.elementType) println(element.text) }}Tests element-based result format with element type assertions on DOCX
using System;using System.Text.Json;using Xberg;
var ConfigOptions = new JsonSerializerOptions { PropertyNameCaseInsensitive = true };var result = await XbergConverter.ExtractAsync(new ExtractInput { Kind = JsonSerializer.Deserialize<ExtractInputKind>("\"uri\"", ConfigOptions)!, Uri = "https://example.com/docx/unit_test_headers.docx" }, new ExtractionConfig { ResultFormat = JsonSerializer.Deserialize<ResultFormat>("\"element_based\"", ConfigOptions)! });foreach (var element in result.Results[0].Elements!){ Console.WriteLine(element.ElementType); Console.WriteLine(element.Text);}Tests element-based result format with element type assertions on DOCX
import Xberg
let result = try await Xberg.extract("{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/unit_test_headers.docx\"}", "{\"result_format\":\"element_based\"}")print(result)Tests element-based result format with element type assertions on DOCX
require "xberg"result = Xberg.extract(Xberg::ExtractInput.new(kind: 'uri', uri: 'https://example.com/docx/unit_test_headers.docx'), { 'result_format' => 'element_based' })(result.results[0].elements || []).each do |element| puts element.element_type.inspect puts element.text.inspectendTests element-based result format with element type assertions on DOCX
<?php
declare(strict_types=1);
require_once __DIR__ . '/vendor/autoload.php';
use Xberg\Xberg;use Xberg\ExtractInput;$input = \Xberg\ExtractInput::from_json(json_encode(["kind" => "uri", "uri" => "https://example.com/docx/unit_test_headers.docx"]));$result = Xberg::extract($input, ["result_format" => "element_based"]);foreach ($result->getResults()[0]->getElements() ?? [] as $element) { var_dump($element->elementType); var_dump($element->text);}Tests element-based result format with element type assertions on DOCX
input_value = %Xberg.ExtractInput{kind: "uri", uri: "https://example.com/docx/unit_test_headers.docx"}result = Xberg.extract_async(input_value, "{\"result_format\":\"element_based\"}")Enum.each(Enum.at(result.results, 0).elements || [], fn element -> IO.inspect(element.element_type) IO.inspect(element.text)end)Tests element-based result format with element type assertions on DOCX
import 'dart:io';import 'package:xberg/xberg.dart';import 'package:xberg/src/xberg_bridge_generated/frb_generated.dart' show RustLib;Future<void> main() async { await RustLib.init(); try { final input = await createExtractInputFromJson(json: '{"kind":"uri","uri":"https://example.com/docx/unit_test_headers.docx"}'); final config = await createExtractionConfigFromJson(json: '{"result_format":"element_based"}'); final result = await XbergBridge.extract(input, config: config); for (final element in result.results[0].elements ?? []) { stdout.writeln(element.elementType); stdout.writeln(element.text); } } finally { RustLib.dispose(); }}Tests element-based result format with element type assertions on DOCX
const std = @import("std");const xberg = @import("xberg");
pub fn main() !void { const _result_json = try xberg.extract("{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/unit_test_headers.docx\"}", "{\"result_format\":\"element_based\"}"); defer std.heap.c_allocator.free(_result_json); std.debug.print("{s}\n", .{_result_json});
}Tests element-based result format with element type assertions on DOCX
#include <assert.h>#include <stdint.h>#include <stdio.h>#include <stdlib.h>#include <string.h>#include "xberg.h"
int main(void) { XBERGAlefHandle input_handle = xberg_extract_input_from_json("{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/unit_test_headers.docx\"}"); XBERGAlefHandle config_handle = xberg_extraction_config_from_json("{\"result_format\":\"element_based\"}"); XBERGAlefHandle result = xberg_extract(input_handle, config_handle); xberg_extract_input_free(input_handle); xberg_extraction_config_free(config_handle); xberg_extraction_result_free(result); return EXIT_SUCCESS;}Elements are in result.elements. Each element has element_id, element_type, text, and metadata.
Element Types
Section titled “Element Types”element_type |
Description | Key additional fields |
|---|---|---|
title |
Main title or top-level heading | level (h1–h6), font_size, font_name |
heading |
Section/subsection heading | level (h1–h6) |
narrative_text |
Body paragraph | — |
list_item |
Bullet, numbered, or indented item | list_type, list_marker, indent_level |
table |
Tabular data | row_count, column_count, format |
image |
Embedded image | format, width, height, alt_text |
code_block |
Code snippet | language, line_count |
block_quote |
Quoted text | — |
header |
Recurring page header | position |
footer |
Recurring page footer | position |
page_break |
Page boundary marker | next_page |
Metadata
Section titled “Metadata”Every element’s metadata contains:
| Field | Type | Description |
|---|---|---|
page_number |
int | None |
1-indexed page number (PDF, DOCX, PPTX) |
filename |
str | None |
Source filename |
coordinates |
BoundingBox | None |
x0, y0, x1, y1 in PDF points. Only populated for text elements when pdf_options.hierarchy is enabled with include_bbox=True. Table and image elements do not carry coordinates. |
element_index |
int |
Zero-based position in the elements array |
additional |
dict[str, str] |
Element-type-specific fields (see table above) |
PDF coordinates use bottom-left origin in points (1/72 inch).
Example Output
Section titled “Example Output”{ "element_id": "elem-a3f2b1c4", "element_type": "title", "text": "Introduction to Machine Learning", "metadata": { "page_number": 1, "element_index": 0, "coordinates": { "x0": 72.0, "y0": 700.0, "x1": 540.0, "y1": 730.0 }, "additional": { "level": "h1", "font_size": "24" } }}Filtering Elements
Section titled “Filtering Elements”config = ExtractionConfig(result_format="element_based")result = extract("document.pdf", config=config)
titles = [e for e in result.elements if e.element_type == "title"]tables = [e for e in result.elements if e.element_type == "table"]
for title in titles: level = title.metadata.additional.get("level", "h1") print(f"[{level}] {title.text}")Migrating from Unstructured.io
Section titled “Migrating from Unstructured.io”If you’re migrating from Unstructured.io, element-based output follows a similar structure with these key differences:
| Aspect | Unstructured.io | Xberg |
|---|---|---|
| Type names | PascalCase (Title, NarrativeText) |
snake_case (title, narrative_text) |
| Element IDs | Not always present | Always present (deterministic hash) |
| Metadata | Basic (page_number, filename) |
Extended (coordinates, additional fields) |
| Config key | — | result_format="element_based" |
Document Structure
Section titled “Document Structure”A flat list of nodes with explicit parent-child index references — a traversable tree with heading levels, content layers, inline annotations, and structured table grids.
Use when you need hierarchical relationships between sections.
Comparison
Section titled “Comparison”| Aspect | Unified (default) | Element-based | Document structure |
|---|---|---|---|
| Output shape | content: string |
elements: array |
nodes: array with index refs |
| Hierarchy | None | Inferred from levels | Explicit parent/child indices |
| Inline annotations | No | No | Bold, italic, links per node |
| Tables | result.tables |
Table elements | TableGrid with cell coords |
| Content layers | Not classified | Not classified | body, header, footer, footnote |
| Best for | LLM prompts, full-text | RAG chunking | Knowledge graphs, structured apps |
Enable
Section titled “Enable”Tests document structure with DOCX heading-driven nesting
import asynciofrom xberg import extract, ExtractInput, ExtractInputKindfrom xberg._xberg import ExtractionConfig
async def main() -> None: input = ExtractInput(kind=ExtractInputKind("uri"), uri="https://example.com/docx/fake.docx") config = ExtractionConfig.from_json("{\"include_document_structure\":true}") result = await extract(input, config) print(result.results[0].document)
asyncio.run(main())Tests document structure with DOCX heading-driven nesting
import { ExtractInput, ExtractInputKind, ExtractionConfig, extract } from "@xberg-io/xberg";async function main() { const input: ExtractInput = { kind: ExtractInputKind.Uri, uri: "https://example.com/docx/fake.docx" }; const config: ExtractionConfig = { includeDocumentStructure: true }; const result = await extract(input, config); console.log(result.results?.[0]?.document);}
void main();Tests document structure with DOCX heading-driven nesting
import { WasmExtractInput, WasmExtractInputKind, extract } from "@xberg-io/xberg-wasm";async function main() { const input: WasmExtractInput = (() => { const _u0 = WasmExtractInput.default(); _u0.kind = WasmExtractInputKind.Uri; _u0.uri = "https://example.com/docx/fake.docx"; return _u0; })(); const result = await extract(input, { includeDocumentStructure: true }); console.log(result.results[0].document);}
void main();Tests document structure with DOCX heading-driven nesting
use xberg::extract;use xberg::ExtractInput;
#[tokio::main]async fn main() { let input_json: serde_json::Value = serde_json::from_str(r#"{"kind":"uri","uri":"https://example.com/docx/fake.docx"}"#).unwrap(); let input = serde_json::from_value::<ExtractInput>(input_json).unwrap(); let config_json: serde_json::Value = serde_json::from_str(r#"{"include_document_structure":true}"#).unwrap(); let config = serde_json::from_value(config_json).unwrap(); let result = extract(input, &config).await.expect("call failed"); println!("{:?}", result.results[0].document);}Tests document structure with DOCX heading-driven nesting
package main
import ( "fmt" xberg "github.com/xberg-io/xberg/packages/go")
func ptr[T any](value T) *T { return &value }func main() { input := xberg.ExtractInput{ Kind: ptr(xberg.ExtractInputKindURI), URI: ptr(`https://example.com/docx/fake.docx`), } config := xberg.ExtractionConfig{ IncludeDocumentStructure: true, } result, err := xberg.Extract(input, config) if err != nil { panic(err) } fmt.Printf("%+v\n", result.Results[0].Document)}Tests document structure with DOCX heading-driven nesting
import io.xberg.*;
public final class Example { public static void main(String[] args) throws Exception { var inputJson = "{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/fake.docx\"}"; var input = JsonUtil.fromJson(inputJson, ExtractInput.class); var configJson = "{\"include_document_structure\":true}"; var config = JsonUtil.fromJson(configJson, ExtractionConfig.class); var result = Xberg.extract(input, config); System.out.println(result.results().get(0).document()); }}Tests document structure with DOCX heading-driven nesting
import io.xberg.*import com.fasterxml.jackson.module.kotlin.jacksonObjectMapper
fun main() = kotlinx.coroutines.runBlocking { val mapper = jacksonObjectMapper().setPropertyNamingStrategy(com.fasterxml.jackson.databind.PropertyNamingStrategies.SNAKE_CASE) val input = mapper.readValue("{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/fake.docx\"}", ExtractInput::class.java) val config = mapper.readValue("{\"include_document_structure\":true,\"url\":{\"crawl\":{\"ssrf\":{}}}}", ExtractionConfig::class.java) val result = Xberg.extract(input, config) println(result.results.first().document)}Tests document structure with DOCX heading-driven nesting
using System;using System.Text.Json;using Xberg;
var ConfigOptions = new JsonSerializerOptions { PropertyNameCaseInsensitive = true };var result = await XbergConverter.ExtractAsync(new ExtractInput { Kind = JsonSerializer.Deserialize<ExtractInputKind>("\"uri\"", ConfigOptions)!, Uri = "https://example.com/docx/fake.docx" }, new ExtractionConfig { IncludeDocumentStructure = true });Console.WriteLine(result.Results[0].Document);Tests document structure with DOCX heading-driven nesting
import Xberg
let result = try await Xberg.extract("{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/fake.docx\"}", "{\"include_document_structure\":true}")debugPrint(result.results()[0].document() as Any)Tests document structure with DOCX heading-driven nesting
require "xberg"result = Xberg.extract(Xberg::ExtractInput.new(kind: 'uri', uri: 'https://example.com/docx/fake.docx'), { 'include_document_structure' => true })puts result.results[0].document.inspectTests document structure with DOCX heading-driven nesting
<?php
declare(strict_types=1);
require_once __DIR__ . '/vendor/autoload.php';
use Xberg\Xberg;use Xberg\ExtractInput;$input = \Xberg\ExtractInput::from_json(json_encode(["kind" => "uri", "uri" => "https://example.com/docx/fake.docx"]));$result = Xberg::extract($input, ["include_document_structure" => true]);var_dump($result->getResults()[0]->getDocument());Tests document structure with DOCX heading-driven nesting
input_value = %Xberg.ExtractInput{kind: "uri", uri: "https://example.com/docx/fake.docx"}result = Xberg.extract_async(input_value, "{\"include_document_structure\":true}")IO.inspect(Enum.at(result.results, 0).document)Tests document structure with DOCX heading-driven nesting
import 'dart:io';import 'package:xberg/xberg.dart';import 'package:xberg/src/xberg_bridge_generated/frb_generated.dart' show RustLib;Future<void> main() async { await RustLib.init(); try { final input = await createExtractInputFromJson(json: '{"kind":"uri","uri":"https://example.com/docx/fake.docx"}'); final config = await createExtractionConfigFromJson(json: '{"include_document_structure":true}'); final result = await XbergBridge.extract(input, config: config); stdout.writeln(result.results[0].document); } finally { RustLib.dispose(); }}Tests document structure with DOCX heading-driven nesting
const std = @import("std");const xberg = @import("xberg");
pub fn main() !void { const _result_json = try xberg.extract("{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/fake.docx\"}", "{\"include_document_structure\":true}"); defer std.heap.c_allocator.free(_result_json); std.debug.print("{s}\n", .{_result_json});
}Tests document structure with DOCX heading-driven nesting
#include <assert.h>#include <stdint.h>#include <stdio.h>#include <stdlib.h>#include <string.h>#include "xberg.h"
int main(void) { XBERGAlefHandle input_handle = xberg_extract_input_from_json("{\"kind\":\"uri\",\"uri\":\"https://example.com/docx/fake.docx\"}"); XBERGAlefHandle config_handle = xberg_extraction_config_from_json("{\"include_document_structure\":true}"); XBERGAlefHandle result = xberg_extract(input_handle, config_handle); xberg_extract_input_free(input_handle); xberg_extraction_config_free(config_handle); xberg_extraction_result_free(result); return EXIT_SUCCESS;}Node Shape
Section titled “Node Shape”Each node in result.document.nodes:
{ "id": "node-a3f2b1c4", "content": { "node_type": "heading", "level": 2, "text": "Supervised Learning" }, "parent": 0, "children": [4, 5, 6], "content_layer": "body", "page": 5, "page_end": null, "bbox": { "x0": 72.0, "y0": 600.0, "x1": 400.0, "y1": 620.0 }, "annotations": []}parentandchildrenare integer indices into thenodesarray (nullif absent)bboxis present when bounding box data is availableannotationscontains inline formatting spans
Node Types
Section titled “Node Types”node_type |
Key fields | Notes |
|---|---|---|
title |
text |
Document title |
heading |
level (1–6), text |
Section heading |
paragraph |
text |
Body paragraph; may have annotations |
list |
ordered (bool) |
Container; children are list_item nodes |
list_item |
text |
Child of list |
table |
grid (TableGrid) |
Grid with cell-level data |
image |
description, image_index |
image_index references result.images |
code |
text, language |
Code block |
quote |
(container) | Children are typically paragraphs |
formula |
text |
Math formula (plain text, LaTeX, or MathML) |
footnote |
text |
Usually content_layer: "footnote" |
group |
label, heading_level, heading_text |
Section grouping container |
page_break |
(marker) | Page boundary |
Content Layers
Section titled “Content Layers”| Layer | Description |
|---|---|
body |
Main document content |
header |
Page header area (repeated chapter titles) |
footer |
Page footer area (page numbers, copyright) |
footnote |
Footnotes and endnotes |
for node in result.document["nodes"]: if node["content_layer"] == "body": process_main_content(node)Text Annotations
Section titled “Text Annotations”Paragraphs carry a list of annotations marking character spans:
{ "start": 0, "end": 16, "kind": { "annotation_type": "bold" } }annotation_type |
Extra fields |
|---|---|
bold, italic, underline, strikethrough |
— |
code, subscript, superscript |
— |
link |
url, title (optional) |
for node in result.document["nodes"]: for ann in node.get("annotations", []): text = node["content"].get("text", "") span = text[ann["start"]:ann["end"]] kind = ann["kind"]["annotation_type"] if kind == "link": print(f"Link: {span} -> {ann['kind']['url']}") else: print(f"{kind}: {span}")Table Grid
Section titled “Table Grid”Table nodes contain a grid with cell-level data:
{ "rows": 3, "cols": 3, "cells": [ { "content": "Algorithm", "row": 0, "col": 0, "row_span": 1, "col_span": 1, "is_header": true }, { "content": "Decision Tree", "row": 1, "col": 0, "row_span": 1, "col_span": 1, "is_header": false } ]}Each cell has row, col, row_span, col_span, is_header, and optionally bbox.
for node in result.document["nodes"]: if node["content"]["node_type"] == "table": grid = node["content"]["grid"] rows, cols = grid["rows"], grid["cols"] table = [[None] * cols for _ in range(rows)] for cell in grid["cells"]: table[cell["row"]][cell["col"]] = cell["content"] for row in table: print(" | ".join(str(c or "") for c in row))PDF Hierarchy Detection
Section titled “PDF Hierarchy Detection”Classifies PDF text blocks into heading levels (H1–H6) and body text via K-means clustering on font sizes — largest cluster is H1, second-largest H2, and so on.
Quick Start
Section titled “Quick Start”import asynciofrom xberg import ExtractInput, extract, ExtractionConfig, PdfConfig, HierarchyConfig
async def main() -> None: config: ExtractionConfig = ExtractionConfig( pdf_options=PdfConfig( extract_metadata=True, hierarchy=HierarchyConfig( enabled=True, k_clusters=6, include_bbox=True, ) ) )
result = await extract(ExtractInput(uri="document.pdf"), config)
# Access hierarchy information for page in result.results[0].pages or []: print(f"Page {page.page_number}:") print(f" Content: {page.content[:100]}...")
asyncio.run(main())import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = { pdfOptions: { extractMetadata: true, hierarchy: { enabled: true, kClusters: 6, includeBbox: true, ocrCoverageThreshold: 0.8, }, },};
const output = await extract({ kind: ExtractInputKind.Uri, uri: "document.pdf" }, config);const result = output.results?.[0];if (result?.pages) { result.pages.forEach((page) => { console.log(`Page ${page.pageNumber}:`); console.log(` Content: ${page.content.substring(0, 100)}...`); });}use xberg::{extract, ExtractionConfig, ExtractInput, PdfConfig, HierarchyConfig};
#[tokio::main]async fn main() -> xberg::Result<()> { let config = ExtractionConfig { pdf_options: Some(PdfConfig { hierarchy: Some(HierarchyConfig { enabled: true, k_clusters: 3, include_bbox: true, }), ..Default::default() }), ..Default::default() };
let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?; for page in output.results[0].pages.iter().flatten() { if let Some(hierarchy) = &page.hierarchy { println!("Page {}: {} hierarchy blocks", page.page_number, hierarchy.block_count); } } Ok(())}package main
import "github.com/xberg-io/xberg/packages/go"
func main() { enabled := true includeBbox := true kClusters := uint(6) kClustersAdvanced := uint(12)
// Basic hierarchy configuration config := xberg.ExtractionConfig{ PdfOptions: &xberg.PdfConfig{ ExtractImages: true, Hierarchy: &xberg.HierarchyConfig{ Enabled: &enabled, KClusters: &kClusters, IncludeBbox: &includeBbox, }, }, }
// Advanced hierarchy configuration with more clusters advancedConfig := xberg.ExtractionConfig{ PdfOptions: &xberg.PdfConfig{ ExtractImages: true, Hierarchy: &xberg.HierarchyConfig{ Enabled: &enabled, KClusters: &kClustersAdvanced, IncludeBbox: &includeBbox, }, }, }
_ = config _ = advancedConfig}import io.xberg.ExtractionConfig;import io.xberg.PdfConfig;import io.xberg.HierarchyConfig;
ExtractionConfig config = ExtractionConfig.builder() .withPdfOptions(PdfConfig.builder() .withHierarchy(HierarchyConfig.builder() .withEnabled(true) .withKClusters(3L) .withIncludeBbox(true) .build()) .build()) .build();using Xberg;
// Basic hierarchy configuration with propertiesvar config = new ExtractionConfig{ PdfOptions = new PdfConfig { ExtractImages = true, Hierarchy = new HierarchyConfig { Enabled = true, KClusters = 6, IncludeBbox = true } }};
var basicResult = (await XbergConverter.ExtractAsync(ExtractInput.FromUri("document.pdf"), config)).Results[0];Console.WriteLine($"Content length: {basicResult.Content.Length}");
// Advanced hierarchy detection with custom parametersvar advancedConfig = new ExtractionConfig{ PdfOptions = new PdfConfig { ExtractImages = true, Hierarchy = new HierarchyConfig { Enabled = true, KClusters = 12, // More clusters for detailed hierarchy IncludeBbox = true // Include bounding box coordinates } }};
var advancedResult = (await XbergConverter.ExtractAsync(ExtractInput.FromUri("complex_document.pdf"), advancedConfig)).Results[0];Console.WriteLine($"Advanced hierarchy detection completed: {advancedResult.Content.Length} chars");
// Minimal configuration with only enabled flagvar minimalConfig = new ExtractionConfig{ PdfOptions = new PdfConfig { Hierarchy = new HierarchyConfig { Enabled = true, // Other properties use defaults: // KClusters = 6 // IncludeBbox = true } }};
var minimalResult = (await XbergConverter.ExtractAsync(ExtractInput.FromUri("document.pdf"), minimalConfig)).Results[0];Console.WriteLine("Extraction with default hierarchy settings complete");
// Disabling hierarchy detectionvar noHierarchyConfig = new ExtractionConfig{ PdfOptions = new PdfConfig { Hierarchy = new HierarchyConfig { Enabled = false } }};
var noHierarchyResult = (await XbergConverter.ExtractAsync(ExtractInput.FromUri("document.pdf"), noHierarchyConfig)).Results[0];Console.WriteLine("Extraction without hierarchy detection complete");require 'xberg'
# Using keyword arguments with defaultsconfig = Xberg::ExtractionConfig.new( pdf_options: Xberg::PdfConfig.new( extract_images: true, hierarchy: Xberg::HierarchyConfig.new( enabled: true, k_clusters: 6, include_bbox: true, ocr_coverage_threshold: 0.8 ) ))
# Using hash syntax alternativeconfig = Xberg::ExtractionConfig.new( pdf_options: Xberg::PdfConfig.new( extract_images: true, hierarchy: { enabled: true, k_clusters: 6, include_bbox: true, ocr_coverage_threshold: 0.8 } ))Output
Section titled “Output”Hierarchy data is in result.pages[n].hierarchy. Each page has a blocks list:
{ "block_count": 4, "blocks": [ { "text": "Chapter 1: Introduction", "level": "h1", "font_size": 24.0, "bbox": [50.0, 100.0, 400.0, 125.0] }, { "text": "Background", "level": "h2", "font_size": 18.0, "bbox": [50.0, 150.0, 300.0, 168.0] }, { "text": "This chapter provides...", "level": "body", "font_size": 12.0, "bbox": [50.0, 200.0, 550.0, 450.0] } ]}bbox:[left, top, right, bottom]in PDF points (present wheninclude_bbox=True). This is the only way to obtain bounding box coordinates for text elements — they are not included by default.level:"h1"–"h6"or"body"
Configuration
Section titled “Configuration”| Parameter | Type | Default | Description |
|---|---|---|---|
enabled |
bool |
true |
Enable hierarchy extraction |
k_clusters |
int |
3 |
Font size clusters (1–7), maps to heading levels |
include_bbox |
bool |
true |
Include bounding box coordinates |
Choosing k_clusters
Section titled “Choosing k_clusters”k_clusters |
Heading levels | Use when |
|---|---|---|
| 2–3 (default) | H1–H2 | Simple documents with 1–2 heading sizes |
| 4–5 | H1–H4 | Standard documents |
| 6–7 | H1–H6+ | Books, specs with deep nesting |
Troubleshooting
Section titled “Troubleshooting”hierarchyisNone— Checkhierarchy.enabledisTrue. This field is only populated by the native PDF text-tree pipeline; it staysNonefor OCR-backed (image-only) pages regardless of whether OCR is enabled — enabling OCR does not populatehierarchy. If fewer text blocks thank_clusters, reducek_clusters. For structured output from a scanned PDF, useoutput_format(e.g.markdown) instead of relying onhierarchy.- Most blocks classified as
body— Document may use uniform font sizes. Reducek_clusters(try 3–4). - Heading levels don’t match visual inspection — Levels are assigned by font size rank, not absolute size. Filter on
block.font_sizedirectly for absolute thresholds.
See the HierarchyConfig reference for the full parameter list.