OCR (Optical Character Recognition)
Extract text from images and scanned PDFs. Xberg automatically determines when OCR is needed — images always require it, scanned PDFs trigger it per-page, and hybrid PDFs only OCR the pages that lack a text layer. Set force_ocr=True to OCR all pages regardless.
See the OcrConfig reference for all configuration options.
Backend Comparison
Section titled “Backend Comparison”Eight OCR backends — pick based on platform, accuracy needs, and language coverage.
All four Candle backends drop an output line that reads as noise rather than recognized text: a line written mostly in a script your configured ocr.language list does not cover (for example, a CJK run on an English-only document), or a plain-text OCR line that is mostly bare LaTeX markup with no ordinary word in it. A $$/``` fenced block, and a line produced by an explicit table/formula/chart task, are never touched. A dropped line adds a ProcessingWarning naming how many lines were removed.
| Tesseract | PaddleOCR | Sceptre | Candle GLM-OCR | Candle TrOCR | Candle DeepSeek-OCR | Candle PaddleOCR-VL | VLM | |
|---|---|---|---|---|---|---|---|---|
| Speed | Fast | Model-dependent | Fast on CPU | Moderate | Moderate | Moderate | Moderate | Slow (API latency) |
| Accuracy | Good | Excellent | Good | Excellent | Good | Excellent | Excellent | Highest |
| Languages | 100+ | 80+ (11 script families) | 8 EasyOCR gen2 groups | All | 100+ | 20+ (CJK + Latin) | 20+ (CJK + Latin) | All (provider-dependent) |
| Installation | System package | Built-in (native) or Python package | Built-in or Cargo feature | Built-in | Built-in | Built-in | Built-in | API key only |
| Model size | ~10 MB | Mobile ~8 MB, Server ~120 MB | Downloaded on first use | ~3 GB | ~250 MB | ~4 GB | ~2.5 GB | None (cloud-hosted) |
| GPU support | No | Yes | No (CPU-only in Xberg 1.1) | Yes (Metal/CUDA) | Yes (Metal/CUDA) | Yes (Metal/CUDA) | Yes (Metal/CUDA) | N/A (server-side) |
| Platform | All (including Wasm) | All except Wasm | Native desktop/server only | Native only | Native only | Native only | Native only | All |
| Cost | Free | Free | Free | Free | Free | Free | Free | Per-token API cost |
Choose a backend
Section titled “Choose a backend”| Backend | Strengths | Trade-offs | Choose it when |
|---|---|---|---|
| Tesseract | Broad language and platform support; low runtime overhead; hierarchical hOCR output | Requires a system package and separately installed language data; less reliable on scene text and difficult scans | You need the default backend, WebAssembly support, or a small deployment footprint |
| PaddleOCR | Strong recognition quality; mobile and server model tiers; GPU support; strong CJK coverage | Larger model downloads; server models can be slow and memory-intensive on CPU | Recognition quality is the priority, especially for CJK documents, and you can budget for the model/runtime cost |
| Sceptre | CRAFT detection and CRNN recognition; eight EasyOCR Gen2 language groups | CPU-only; tract mobile builds and opt-in worker-hosted WASM have larger artifacts | You need EasyOCR Gen2 recognition without a Python runtime |
The remaining backends cover larger-model and hosted use cases:
- Candle GLM-OCR — Excellent accuracy with VLM-level reasoning on 0.9B-param GLM model. Pure Rust, GPU-accelerated (Metal on macOS, CUDA on Linux). Region-aware layout dispatch. First download ~3 GB.
- Candle TrOCR — Smaller model footprint (~250 MB) with solid accuracy across languages. Pure Rust, GPU-accelerated. Good balance of speed and quality.
- Candle DeepSeek-OCR — Deep learning-based OCR combining SAM + CLIP + Qwen2 + DeepSeek MoE. Multilingual with strong CJK coverage. Pure Rust, GPU-accelerated. First download 6.7 GB (BF16 weights).
- Candle PaddleOCR-VL — SigLIP vision encoder + Ernie-4.5 text decoder. Lightweight multilingual model with CJK and Latin support. Pure Rust, GPU-accelerated. First download ~2.5 GB.
- VLM — Best for handwritten text, poor scans, Arabic/Farsi, and complex layouts. Requires an API key and incurs per-token costs. See LLM Integration for full details.
Installation
Section titled “Installation”Tesseract
Section titled “Tesseract”brew install tesseractsudo apt-get install tesseract-ocrsudo dnf install tesseractDownload from GitHub releases.
Additional language packs:
# macOS — all languagesbrew install tesseract-lang
# Ubuntu/Debian — individual languagessudo apt-get install tesseract-ocr-deu # Germansudo apt-get install tesseract-ocr-fra # French
# Verify installed languagestesseract --list-langsTesseract in Wasm
Section titled “Tesseract in Wasm”The Wasm package has no packaged helper that wires up a Tesseract backend for you (there is no
enableOcr() export). The primitive it does export is registerOcrBackend() — you supply a
JavaScript object implementing the OcrBackend trait bridge and register it under the name
"tesseract-wasm" before calling extract with { ocr: { backend: "tesseract-wasm" } }. See
Trait Bridges in the WASM API Reference for the registration
API.
Verify Tesseract from your binding
Section titled “Verify Tesseract from your binding”Run one OCR extraction to confirm that your binding finds Tesseract. force_ocr makes the pipeline run OCR even on a PDF that already has a text layer:
#include <xberg.h>#include <stdio.h>#include <stdlib.h>#include <string.h>
int main(void) { const char *config_json = "{" "\"force_ocr\": true," "\"ocr\": {\"backend\": \"tesseract\", \"language\": \"eng\"}" "}";
XBERGAlefHandle config = xberg_extraction_config_from_json(config_json); if (config == 0) { fprintf(stderr, "config init failed (code %d): %s\n", xberg_last_error_code(), xberg_last_error_context()); return 1; }
XBERGAlefHandle input = xberg_extract_input_from_uri("scanned.pdf"); if (input == 0) { fprintf(stderr, "Failed to create input (code %d): %s\n", xberg_last_error_code(), xberg_last_error_context()); xberg_extraction_config_free(config); return 1; }
XBERGAlefHandle result = xberg_extract(input, config); if (result == 0) { fprintf(stderr, "extraction failed (code %d): %s\n", xberg_last_error_code(), xberg_last_error_context()); xberg_extract_input_free(input); xberg_extraction_config_free(config); return 1; }
char *content = xberg_extraction_result_results(result); printf("%s\n", content ? content : "(empty)"); xberg_free_string(content);
xberg_extract_input_free(input); xberg_extraction_result_free(result); xberg_extraction_config_free(config); return 0;}using Xberg;
var config = new ExtractionConfig{ Ocr = new OcrConfig { Backend = "tesseract", Language = ["eng", "deu", "fra"], TesseractConfig = new TesseractConfig { Psm = 3 } }};
var result = (await XbergConverter.ExtractAsync(ExtractInput.FromUri("document.pdf"), config)).Results[0];Console.WriteLine(result.Content);package main
import ( "fmt" "log"
"github.com/xberg-io/xberg/packages/go")
func main() { ocrConfig := &xberg.OcrConfig{ Backend: xberg.Ptr("tesseract"), Language: []string{"eng"}, }
config := xberg.ExtractionConfig{ Ocr: ocrConfig, }
input := xberg.ExtractInputFromURI("scanned.pdf") result, err := xberg.Extract(*input, config) if err != nil { log.Fatalf("extract failed: %v", err) }
fmt.Println("Extracted text from scanned document:") fmt.Println(result.Results[0].Content) fmt.Println("Used OCR backend: tesseract")}<?php
declare(strict_types=1);
/** * Basic OCR with Tesseract * * Extract text from scanned PDFs and images using Tesseract OCR. */
require_once __DIR__ . '/vendor/autoload.php';
use Xberg\ExtractionConfig;use Xberg\OcrConfig;
$config = new ExtractionConfig( ocr: new OcrConfig( backend: 'tesseract', language: 'eng' ));
$output = \Xberg\XbergApi::extract(\Xberg\ExtractInput::fromUri('scanned_document.pdf'), $config ?? \Xberg\ExtractionConfig::default());$result = $output->getResults()[0];
echo "OCR Extraction Results:\n";echo str_repeat('=', 60) . "\n";echo $result->content . "\n\n";
$multilingualConfig = new ExtractionConfig( ocr: new OcrConfig( backend: 'tesseract', language: 'eng+fra+deu' ));
$output = \Xberg\XbergApi::extract(\Xberg\ExtractInput::fromUri('multilingual_scan.pdf'), $config ?? \Xberg\ExtractionConfig::default());$result = $output->getResults()[0];
echo "Multilingual OCR:\n";echo str_repeat('=', 60) . "\n";echo substr($result->content, 0, 500) . "...\n\n";
$imageConfig = new ExtractionConfig( ocr: new OcrConfig( backend: 'tesseract', language: 'eng' ));
$imageFormats = ['png', 'jpg', 'tiff'];foreach ($imageFormats as $format) { $file = "scan.$format"; if (file_exists($file)) { echo "Processing $file...\n"; $output = \Xberg\XbergApi::extract(\Xberg\ExtractInput::fromUri($file), $config ?? \Xberg\ExtractionConfig::default());$result = $output->getResults()[0]; echo "Extracted " . strlen($result->content) . " characters\n"; echo "Preview: " . substr($result->content, 0, 100) . "...\n\n"; }}
$languages = [ 'spa' => 'Spanish document', 'fra' => 'French document', 'deu' => 'German document', 'ita' => 'Italian document', 'por' => 'Portuguese document', 'rus' => 'Russian document', 'jpn' => 'Japanese document', 'chi_sim' => 'Chinese (Simplified) document',];
foreach ($languages as $lang => $description) { $file = strtolower(str_replace(' ', '_', $description)) . '.pdf';
if (file_exists($file)) { $config = new ExtractionConfig( ocr: new OcrConfig( backend: 'tesseract', language: $lang ) );
$output = \Xberg\XbergApi::extract(\Xberg\ExtractInput::fromUri($file), $config ?? \Xberg\ExtractionConfig::default());$result = $output->getResults()[0];
echo "$description ($lang):\n"; echo " Characters extracted: " . mb_strlen($result->content) . "\n\n"; }}
$config = new ExtractionConfig( ocr: new OcrConfig(backend: 'tesseract', language: 'eng'));
$output = \Xberg\XbergApi::extract(\Xberg\ExtractInput::fromUri('invoice_scan.pdf'), $config);$result = $output->getResults()[0];
echo "Invoice OCR:\n";echo str_repeat('=', 60) . "\n";echo $result->content . "\n";
$output = \Xberg\XbergApi::extract(\Xberg\ExtractInput::fromUri('scanned.pdf'), $config ?? \Xberg\ExtractionConfig::default());$result = $output->getResults()[0];
$contentLength = strlen($result->content);$pageCount = $result->metadata?->pdf?->page_count ?? 1;$avgCharsPerPage = $contentLength / $pageCount;
echo "\nOCR Quality Assessment:\n";echo "Total characters: $contentLength\n";echo "Pages: $pageCount\n";echo "Average chars/page: " . number_format($avgCharsPerPage) . "\n";
if ($avgCharsPerPage < 100) { echo "Warning: Low character count may indicate poor scan quality\n"; echo "Consider using image preprocessing or higher DPI settings.\n";} elseif ($avgCharsPerPage > 2000) { echo "Pass: Good - Adequate text extracted\n";} else { echo "Pass: Moderate - Text extracted successfully\n";}import asynciofrom xberg import ExtractInput, extract, ExtractionConfig, OcrConfig, TesseractConfig
async def main() -> None: config = ExtractionConfig( force_ocr=True, ocr=OcrConfig( backend="tesseract", language="eng", tesseract_config=TesseractConfig(psm=3) ) ) result = await extract(ExtractInput(uri="scanned.pdf"), config) print(result.results[0].content) print(f"Detected Languages: {result.results[0].detected_languages}")
asyncio.run(main())require 'xberg'
ocr_config = Xberg::OcrConfig.new( backend: 'tesseract', language: 'eng')
config = Xberg::ExtractionConfig.new(ocr: ocr_config)input = Xberg::ExtractInput.new(uri: 'scanned.pdf')result = Xberg.extract(input, config)
puts "Extracted text from scanned document:"puts result.results.first.contentputs "Used OCR backend: tesseract"use xberg::{extract, ExtractionConfig, ExtractInput, OcrConfig};
#[tokio::main]async fn main() -> xberg::Result<()> { let config = ExtractionConfig { force_ocr: true, ocr: Some(OcrConfig { backend: "tesseract".to_string(), language: vec!["eng".to_string()], ..Default::default() }), ..Default::default() };
let output = extract(ExtractInput::from_uri("scanned.pdf"), &config).await?; println!("{}", output.results[0].content); println!("Detected languages: {:?}", output.results[0].detected_languages); Ok(())}PaddleOCR
Section titled “PaddleOCR”Built in via the paddle-ocr feature flag. Models download automatically on first use — no extra installation needed.
[dependencies]xberg = { version = "1", features = ["paddle-ocr"] }PaddleOCR is bundled via the native Rust bindings and works out of the box since 4.8.5 — no extra installation is needed. Models are downloaded automatically on first use.
PaddleOCR (tract backend)
Section titled “PaddleOCR (tract backend)”ONNX Runtime cannot link on wasm32 or the Android x86_64 emulator. On those targets, enable paddle-ocr-tract instead of
paddle-ocr-ort (or the paddle-ocr alias) to run PaddleOCR through the pure-Rust tract engine, with no ONNX Runtime in
the dependency graph. It runs the same DBNet detector, CRNN recognizer, and PP-LCNet text-line orientation classifier.
paddle-ocr-tract is included in android-target. See Pure-Rust Inference (tract)
for how the engine handles PaddleOCR’s shape requirements.
[dependencies]xberg = { version = "1", features = ["paddle-ocr-tract"] }Do not enable both an ORT PaddleOCR feature and paddle-ocr-tract in the same build — as with auto-rotate /
auto-rotate-tract and layout-detection / layout-tract, the ORT and tract variants of a feature are
mutually exclusive.
Limitations on tract:
- CPU-only. GPU/execution-provider settings apply to the ONNX Runtime path only and are ignored on tract.
- Slower than ONNX Runtime. tract trades throughput for portability on targets where ONNX Runtime cannot link at all; see Latency for measured ratios on other models.
- Detection builds a plan per page shape. DBNet’s shape cannot be left symbolic under tract, so its plan is pinned to the exact dimensions each page resizes to and cached (four plans, least-recently-used eviction). Pages of a document nearly always share one extent, so the plan is built once; a new extent costs one build. Detection results are identical to the ONNX Runtime path — see PaddleOCR under tract for why the page is never padded into a fixed canvas.
Detection cost grows steeply with page size. Measured on macOS arm64 (single run, indicative, not a benchmark):
| Model | 640² | 960² | 1280² |
|---|---|---|---|
| v6 det medium | 821 ms / 518 MiB | 1650 ms / 981 MiB | 4038 ms / 1652 MiB |
| v2 det mobile | 140 ms / 90 MiB | 224 ms / 175 MiB | 706 ms / 306 MiB |
det_limit_side_len defaults to 1024, at which the medium detector extrapolates to roughly 2 s and ~1.1 GiB per page
under tract. On tract targets, prefer the mobile detection tier and lower det_limit_side_len — around 640 — to keep
per-page latency and memory in check. The ORT default is unchanged.
Sceptre
Section titled “Sceptre”Sceptre is included in published native desktop/server packages and the CLI. These builds use ONNX Runtime and
download checksum-pinned models to the Hugging Face cache on first use. For a custom desktop/server Rust build,
enable sceptre-ocr; add auto-rotate if you also need automatic orientation correction.
[dependencies]xberg = { version = "1", features = ["sceptre-ocr"] }Android and iOS builds use the pure-Rust tract engine through sceptre-ocr-tract; use auto-rotate-tract for
orientation correction. Mobile builds exclude Sceptre’s Hugging Face downloader. Resolve packaged application assets
to accessible filesystem paths and set both backend_options.model.detector_path and
backend_options.model.recognizer_path. WebAssembly support is excluded from the published default bundle because of
model and runtime size. Source builds can enable the separate byte-fed class with
wasm-pack build crates/xberg-wasm --target web --features sceptre-wasm; application code must instantiate and call
that synchronous class inside its own Web Worker. The JavaScript host fetches the CRAFT and recognizer models and
supplies their bytes.
All Sceptre paths are CPU-only.
Configuration
Section titled “Configuration”Basic OCR
Section titled “Basic OCR”import asynciofrom xberg import ExtractInput, extract, ExtractionConfig, OcrConfig
async def main() -> None: config: ExtractionConfig = ExtractionConfig( ocr=OcrConfig(backend="tesseract", language=["eng"]) )
result = await extract(ExtractInput(uri="scanned.pdf"), config)
content: str = result.results[0].content preview: str = content[:100] total_length: int = len(content)
print(f"Extracted content (preview): {preview}") print(f"Total characters: {total_length}")
asyncio.run(main())import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = { ocr: { backend: "tesseract", language: ["eng"], },};
const output = await extract({ kind: ExtractInputKind.Uri, uri: "scanned.pdf" }, config);console.log(output.results?.[0]?.content);use xberg::{extract, ExtractionConfig, ExtractInput, OcrConfig};
#[tokio::main]async fn main() -> xberg::Result<()> { let config = ExtractionConfig { ocr: Some(OcrConfig { backend: "tesseract".to_string(), language: vec!["eng".to_string()], ..Default::default() }), ..Default::default() };
let output = extract(ExtractInput::from_uri("scanned.pdf"), &config).await?; println!("{}", output.results[0].content); Ok(())}package main
import ( "log"
"github.com/xberg-io/xberg/packages/go")
func main() { cfg := xberg.ExtractionConfig{ Ocr: &xberg.OcrConfig{ Backend: xberg.Ptr("tesseract"), Language: []string{"eng"}, }, }
input := xberg.ExtractInputFromURI("scanned.pdf") result, err := xberg.Extract(*input, cfg) if err != nil { log.Fatalf("extract failed: %v", err) } log.Println(len(result.Results[0].Content))}import io.xberg.ExtractInput;import io.xberg.ExtractInputKind;import io.xberg.ExtractedDocument;import io.xberg.ExtractionConfig;import io.xberg.ExtractionResult;import io.xberg.OcrConfig;import io.xberg.Xberg;import io.xberg.XbergRsException;import java.util.List;
public class Main { public static void main(String[] args) { try { ExtractionConfig config = ExtractionConfig.builder() .withForceOcr(true) .withOcr(OcrConfig.builder() .withBackend("tesseract") .withLanguage(List.of("eng")) .build()) .build();
ExtractInput input = ExtractInput.builder() .withKind(ExtractInputKind.URI) .withUri("scanned.pdf") .build();
ExtractionResult output = Xberg.extract(input, config); ExtractedDocument document = output.results().get(0); System.out.println(document.content()); } catch (XbergRsException e) { System.err.println("Extraction failed: " + e.getMessage()); } }}require 'xberg'
ocr_config = Xberg::OcrConfig.new( backend: 'tesseract', language: 'eng')
config = Xberg::ExtractionConfig.new(ocr: ocr_config)input = Xberg::ExtractInput.new(uri: 'scanned.pdf')result = Xberg.extract(input, config)puts result.results.first.contentimport init, { extract } from "@xberg-io/xberg-wasm";
await init();
const fileInput = document.getElementById("file") as HTMLInputElement;const file = fileInput.files?.[0];
if (file) { const bytes = new Uint8Array(await file.arrayBuffer()); const result = await extract( { kind: "bytes", bytes, mimeType: file.type }, { ocr: { enabled: true, backend: "tesseract", language: ["eng"], }, }, ); console.log(result.results[0].content);}import init, { extract } from "@xberg-io/xberg-wasm";
// Outside the browser the default `fetch`-based init cannot read a `file://`// URL: pass the `xberg_wasm_bg.wasm` bytes yourself, either as// `init({ module_or_path: bytes })` or via the synchronous `initSync({ module: bytes })`.await init();
const result = await extract( { kind: "uri", uri: "./scanned_document.png" }, { ocr: { enabled: true, backend: "tesseract", language: ["eng"], }, },);console.log(result.results[0].content);Multiple Languages
Section titled “Multiple Languages”Specify multiple language codes separated by + (Tesseract) or as a list (PaddleOCR and VLM backends):
import asynciofrom xberg import ExtractInput, extract, ExtractionConfig, OcrConfig
async def main() -> None: config: ExtractionConfig = ExtractionConfig( ocr=OcrConfig(backend="tesseract", language=["eng", "deu", "fra"]) )
result = await extract(ExtractInput(uri="multilingual.pdf"), config)
content: str = result.results[0].content preview: str = content[:100] total_length: int = len(content)
print(f"Extracted content (preview): {preview}") print(f"Total characters: {total_length}")
asyncio.run(main())import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = { ocr: { backend: "tesseract", language: ["eng", "deu", "fra"], },};
const output = await extract({ kind: ExtractInputKind.Uri, uri: "multilingual.pdf" }, config);console.log(output.results?.[0]?.content);use xberg::{extract, ExtractionConfig, ExtractInput, OcrConfig};
#[tokio::main]async fn main() -> xberg::Result<()> { let config = ExtractionConfig { ocr: Some(OcrConfig { backend: "tesseract".to_string(), language: vec!["eng".to_string(), "deu".to_string(), "fra".to_string()], ..Default::default() }), ..Default::default() };
let output = extract(ExtractInput::from_uri("multilingual.pdf"), &config).await?; println!("{}", output.results[0].content); Ok(())}package main
import ( "log"
"github.com/xberg-io/xberg/packages/go")
func main() { config := xberg.ExtractionConfig{ Ocr: &xberg.OcrConfig{ Backend: xberg.Ptr("tesseract"), Language: []string{"eng", "deu", "fra"}, }, } input := xberg.ExtractInputFromURI("multilingual.pdf") result, err := xberg.Extract(*input, config) if err != nil { log.Fatalf("extract failed: %v", err) }
log.Println(result.Results[0].Content)}import io.xberg.Xberg;import io.xberg.ExtractInputKind;import io.xberg.ExtractionResult;import io.xberg.ExtractedDocument;import io.xberg.ExtractionConfig;import io.xberg.ExtractInput;import io.xberg.OcrConfig;import java.util.List;
ExtractionConfig config = ExtractionConfig.builder() .withOcr(OcrConfig.builder() .withBackend("tesseract") .withLanguage(List.of("eng", "deu", "fra")) .build()) .build();ExtractionResult output = Xberg.extract( ExtractInput.builder().withKind(ExtractInputKind.URI).withUri("multilingual.pdf").build(), config);ExtractedDocument result = output.results().get(0);System.out.println(result.content());require 'xberg'
config = Xberg::ExtractionConfig.new( ocr: Xberg::OcrConfig.new( backend: 'tesseract', language: 'eng+deu+fra' ))
input = Xberg::ExtractInput.new(uri: 'multilingual.pdf')result = Xberg.extract(input, config)puts result.results.first.content// NOTE: The Wasm package does not ship a ready-made Tesseract adapter or an// enableOcr() helper. Before calling extract, register your own OcrBackend// trait bridge object (name/supportedLanguages/processImage) under the name// "tesseract-wasm" via registerOcrBackend() — see the WASM API Reference's// Trait Bridges section.import init, { ExtractInputKind, extract, registerOcrBackend } from '@xberg-io/xberg-wasm';
await init();registerOcrBackend(myTesseractWasmBackend);
const file = fileInput.files?.[0];if (file) { const output = await extract( { kind: ExtractInputKind.Bytes, bytes: new Uint8Array(await file.arrayBuffer()), mimeType: file.type || 'application/octet-stream', filename: file.name, }, { ocr: { backend: 'tesseract-wasm', language: ['eng', 'deu'] } }, ); const result = output.results[0];}Force OCR
Section titled “Force OCR”Process PDFs with OCR even when they have a text layer:
import asynciofrom xberg import ExtractInput, extract, ExtractionConfig, OcrConfig
async def main() -> None: config: ExtractionConfig = ExtractionConfig( ocr=OcrConfig(backend="tesseract"), force_ocr=True, )
result = await extract(ExtractInput(uri="document.pdf"), config)
content: str = result.results[0].content preview: str = content[:100] total_length: int = len(content)
print(f"Extracted content (preview): {preview}") print(f"Total characters: {total_length}")
asyncio.run(main())import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = { ocr: { backend: "tesseract", }, forceOcr: true,};
const output = await extract({ kind: ExtractInputKind.Uri, uri: "document.pdf" }, config);console.log(output.results?.[0]?.content);use xberg::{extract, ExtractionConfig, ExtractInput, OcrConfig};
#[tokio::main]async fn main() -> xberg::Result<()> { let config = ExtractionConfig { ocr: Some(OcrConfig { backend: "tesseract".to_string(), ..Default::default() }), force_ocr: true, ..Default::default() };
let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?; println!("{}", output.results[0].content); Ok(())}package main
import ( "fmt" "log"
"github.com/xberg-io/xberg/packages/go")
func main() { config := xberg.ExtractionConfig{ Ocr: &xberg.OcrConfig{ Backend: xberg.Ptr("tesseract"), }, ForceOcr: true, } input := xberg.ExtractInputFromURI("document.pdf") result, err := xberg.Extract(*input, config) if err != nil { log.Fatalf("extract failed: %v", err) }
fmt.Println(result.Results[0].Content)}import io.xberg.Xberg;import io.xberg.ExtractInputKind;import io.xberg.ExtractionResult;import io.xberg.ExtractedDocument;import io.xberg.ExtractionConfig;import io.xberg.ExtractInput;import io.xberg.OcrConfig;
ExtractionConfig config = ExtractionConfig.builder() .withOcr(OcrConfig.builder() .withBackend("tesseract") .build()) .withForceOcr(true) .build();ExtractionResult output = Xberg.extract( ExtractInput.builder().withKind(ExtractInputKind.URI).withUri("document.pdf").build(), config);ExtractedDocument result = output.results().get(0);System.out.println(result.content());require 'xberg'
config = Xberg::ExtractionConfig.new( ocr: Xberg::OcrConfig.new(backend: 'tesseract'), force_ocr: true)
input = Xberg::ExtractInput.new(uri: 'document.pdf')result = Xberg.extract(input, config)puts result.results.first.contentDisable OCR
Section titled “Disable OCR”When disable_ocr is set, image files return empty content instead of raising MissingDependencyError:
from xberg import ExtractInput, ExtractionConfig, extract
config = ExtractionConfig(disable_ocr=True)output = await extract(ExtractInput(kind="uri", uri="scanned.png"), config=config)result = output.results[0]# result.content will be empty — OCR was skippedimport { ExtractInputKind, extract } from '@xberg-io/xberg';
const output = await extract( { kind: ExtractInputKind.Uri, uri: 'scanned.png' }, { disableOcr: true },);const result = output.results[0];// result.content will be empty — OCR was skippeduse xberg::{extract, ExtractInput, ExtractionConfig};
let config = ExtractionConfig { disable_ocr: true, ..Default::default()};let output = extract(ExtractInput::from_uri("scanned.png"), &config).await?;let result = &output.results[0];// result.content will be empty — OCR was skippedUsing PaddleOCR
Section titled “Using PaddleOCR”import asynciofrom xberg import ExtractInput, extract, ExtractionConfig, OcrConfig
async def main() -> None: config: ExtractionConfig = ExtractionConfig( ocr=OcrConfig(backend="paddleocr", language=["en"]) # model_tier="server" for max accuracy )
result = await extract(ExtractInput(uri="scanned.pdf"), config)
content: str = result.results[0].content preview: str = content[:100] total_length: int = len(content)
print(f"Extracted content (preview): {preview}") print(f"Total characters: {total_length}")
asyncio.run(main())import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = { ocr: { backend: "paddle-ocr", language: ["en"], // modelTier: 'server', // for max accuracy },};
const output = await extract({ kind: ExtractInputKind.Uri, uri: "scanned.pdf" }, config);console.log(output.results?.[0]?.content);use xberg::{extract, ExtractionConfig, ExtractInput, OcrConfig};
#[tokio::main]async fn main() -> xberg::Result<()> { let config = ExtractionConfig { ocr: Some(OcrConfig { backend: "paddleocr".to_string(), language: vec!["en".to_string()], // paddle_ocr_config: Some(serde_json::json!({"model_tier": "server"})), // for max accuracy ..Default::default() }), ..Default::default() };
let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?; println!("Extracted text: {}", output.results[0].content); Ok(())}package main
import ( "log"
"github.com/xberg-io/xberg/packages/go")
func main() { cfg := xberg.ExtractionConfig{ Ocr: &xberg.OcrConfig{ Backend: xberg.Ptr("paddle-ocr"), Language: []string{"en"}, }, }
input := xberg.ExtractInputFromURI("scanned.pdf") result, err := xberg.Extract(*input, cfg) if err != nil { log.Fatalf("extract failed: %v", err) } log.Println(len(result.Results[0].Content))}import io.xberg.Xberg;import io.xberg.ExtractInputKind;import io.xberg.ExtractionResult;import io.xberg.ExtractedDocument;import io.xberg.XbergRsException;import io.xberg.ExtractionConfig;import io.xberg.ExtractInput;import io.xberg.OcrConfig;import java.util.List;
public class Main { public static void main(String[] args) { try { ExtractionConfig config = ExtractionConfig.builder() .withOcr(OcrConfig.builder() .withBackend("paddle-ocr") .withLanguage(List.of("en")) // .withPaddleOcrConfig(PaddleOcrConfig.builder().withModelTier("server").build()) // for max accuracy .build()) .build(); ExtractionResult output = Xberg.extract( ExtractInput.builder().withKind(ExtractInputKind.URI).withUri("scanned.pdf").build(), config ); ExtractedDocument result = output.results().get(0); System.out.println(result.content()); } catch (XbergRsException e) { System.err.println("Extraction failed: " + e.getMessage()); } }}require 'xberg'
config = Xberg::ExtractionConfig.new( ocr: Xberg::OcrConfig.new( backend: 'paddleocr', language: 'eng' # model_tier: 'server' # for max accuracy ))
input = Xberg::ExtractInput.new(uri: 'scanned.pdf')result = Xberg.extract(input, config)puts result.results.first.content[0..100]puts "Total length: #{result.results.first.content.length}"Using Sceptre
Section titled “Using Sceptre”Select the backend with backend = "sceptre". Put Sceptre’s detection, recognition, concurrency, and model
sections directly under backend_options; do not add a nested sceptre key.
[ocr]backend = "sceptre"language = ["eng", "deu"]
[ocr.backend_options.recognition]batch_size = 4
[ocr.backend_options.concurrency]max_threads = 2| Path | Purpose |
|---|---|
backend_options.detection.* |
CRAFT thresholds, canvas size, magnification, box size, and line grouping |
backend_options.recognition.* |
Decoder, beam width, batch size, allow/block lists, contrast retry, and confidence filtering |
backend_options.concurrency.max_threads |
Cap Sceptre’s worker and inference thread pools |
backend_options.model.cache_dir |
Override the Hugging Face cache root on desktop/server ORT builds |
backend_options.model.registry_owner |
Use an approved mirror of the checksum-pinned model repositories |
backend_options.model.detector_path |
Required CRAFT model path on Android/iOS tract builds |
backend_options.model.recognizer_path |
Required selected Gen2 recognizer path on Android/iOS tract builds |
Set languages with ocr.language; Xberg overrides backend_options.model.languages. Xberg also selects ORT or tract
for the build target, so portable configuration should not set backend_options.model.backend.
[ocr]backend = "sceptre"language = ["deu"]
[ocr.backend_options.model]detector_path = "/app/models/craft_mlt_25k.onnx"recognizer_path = "/app/models/latin_g2.onnx"The paths must be resolved by the application from its bundle or asset system before extraction starts.
Sceptre provides all eight EasyOCR Gen2 recognition groups: english, latin, chinese_simplified, japanese,
korean, cyrillic, telugu, and kannada. Xberg accepts group tokens and ISO aliases such as en/eng,
de/deu, zh/zho, ja/jpn, ko/kor, ru/rus, te/tel, and kn/kan. English may be combined
with one other group; two distinct non-English groups require different recognizers and fail validation before OCR.
Sceptre emits line-level quadrilaterals and recognition confidence. It does not provide Tesseract’s word/symbol hierarchy or a separate detection-confidence value. Tract inference uses a fixed CRAFT canvas and can require more memory than the default browser OCR path; run WebAssembly inference in a worker.
Candle GLM-OCR
Section titled “Candle GLM-OCR”Ships in the published package by default on native desktop/server platforms — no feature flag or custom build needed. Set ocr.backend = "candle-glm-ocr" and the GLM-OCR model downloads automatically on first use (~3 GB), cached at ~/.cache/huggingface/. Not available on WebAssembly, Android, iOS, Dart, or Swift.
For a custom Rust crate build, the model is gated behind the candle-glm-ocr feature (or the candle-vlm-ocr umbrella feature):
[dependencies]xberg = { version = "1", features = ["candle-glm-ocr"] }GPU support:
- Metal (macOS) — Default, F32 dtype (BF16 matmul unavailable in candle 0.10)
- CUDA (Linux/Windows with NVIDIA GPU) — Auto-detected
- CPU fallback — Slowest, but always available
Using Candle GLM-OCR
Section titled “Using Candle GLM-OCR”Candle GLM-OCR dispatches by detected layout region using PP-DocLayout-V3. Each region runs through the appropriate task prompt (ocr/table/formula/chart/caption) and outputs are merged into reading-order markdown.
from xberg import ExtractInput, ExtractionConfig, OcrConfig, extract
# Paired mode: per-region dispatch (default)config = ExtractionConfig( force_ocr=True, ocr=OcrConfig( backend="candle-glm-ocr", language=["en"], backend_options={"layout_mode": "paired"}, ),)output = await extract(ExtractInput(kind="uri", uri="document.pdf"), config=config)result = output.results[0]print(result.content)
# Whole-page mode: single OCR pass over entire pageconfig_whole = ExtractionConfig( force_ocr=True, ocr=OcrConfig( backend="candle-glm-ocr", language=["en"], backend_options={"layout_mode": "whole_page"}, ),)whole_page_output = await extract( ExtractInput(kind="uri", uri="document.pdf"), config=config_whole,)result_whole_page = whole_page_output.results[0]import { ExtractInputKind, extract } from '@xberg-io/xberg';
// Paired mode: per-region dispatch (default)const output = await extract( { kind: ExtractInputKind.Uri, uri: 'document.pdf' }, { forceOcr: true, ocr: { backend: 'candle-glm-ocr', language: ['en'], backendOptions: { layout_mode: 'paired' }, }, },);const result = output.results[0];console.log(result.content);
// Whole-page modeconst wholePageOutput = await extract( { kind: ExtractInputKind.Uri, uri: 'document.pdf' }, { forceOcr: true, ocr: { backend: 'candle-glm-ocr', language: ['en'], backendOptions: { layout_mode: 'whole_page' }, }, },);const resultWholePage = wholePageOutput.results[0];use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig};use serde_json::json;
// Paired mode: per-region dispatch (default)let config = ExtractionConfig { force_ocr: true, ocr: Some(OcrConfig { backend: "candle-glm-ocr".into(), language: vec!["en".to_string()], backend_options: Some(json!({"layout_mode": "paired"})), ..Default::default() }), ..Default::default()};let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?;let result = &output.results[0];println!("{}", result.content);
// Whole-page modelet config_whole = ExtractionConfig { force_ocr: true, ocr: Some(OcrConfig { backend: "candle-glm-ocr".into(), language: vec!["en".to_string()], backend_options: Some(json!({"layout_mode": "whole_page"})), ..Default::default() }), ..Default::default()};let whole_page_output = extract(ExtractInput::from_uri("document.pdf"), &config_whole).await?;let _result_whole_page = &whole_page_output.results[0];Backend options:
| Option | Values | Description |
|---|---|---|
layout_mode |
"paired" (default), "whole_page" |
Paired: dispatch per-region via PP-DocLayout-V3. Whole-page: single OCR pass on entire page. |
task |
"ocr" (default), "table", "formula", "chart", "caption" |
Task prompt for whole-page mode only; ignored in paired mode where the region type determines the prompt. |
device |
"auto" (default), "cpu", "metal", "cuda" |
Device selection. Auto detects Metal on macOS, CUDA on Linux, CPU fallback. |
Candle DeepSeek-OCR
Section titled “Candle DeepSeek-OCR”DeepSeek-OCR — combination of SAM + CLIP encoder fused with Qwen2 decoder and DeepSeek V2 MoE for comprehensive multilingual document understanding. Markdown output.
Ships in the published package by default on native desktop/server platforms — no feature flag or custom build needed. Set ocr.backend = "candle-deepseek-ocr" and the weights download automatically on first use from a checksum-pinned revision of deepseek-ai/DeepSeek-OCR (6.7 GB, BF16) into the standard Hugging Face cache (HF_HUB_CACHE / HUGGINGFACE_HUB_CACHE / HF_HOME, or ~/.cache/huggingface/ by default) — no backend_options.model_path needed. Not available on WebAssembly, Android, iOS, Dart, or Swift.
For a custom Rust crate build, the model is gated behind the candle-deepseek-ocr feature (or the candle-vlm-ocr umbrella feature):
[dependencies]xberg = { version = "1", features = ["candle-deepseek-ocr"] }GPU support: dtype follows the compute device (backend_options.dtype = "auto" by default; a wrong dtype fails the load hard rather than silently falling back, so override it only when you know the target device’s kernel coverage).
- CUDA (Linux/Windows with NVIDIA GPU) — Auto-detected, BF16
- Metal (macOS) — F16 (BF16 kernel coverage is incomplete on Metal)
- CPU fallback — F32, slowest but always available
Using Candle DeepSeek-OCR
Section titled “Using Candle DeepSeek-OCR”from xberg import ExtractInput, ExtractionConfig, OcrConfig, extract
config = ExtractionConfig( force_ocr=True, ocr=OcrConfig( backend="candle-deepseek-ocr", language=["en"], backend_options={"device": "auto"}, ),)output = await extract(ExtractInput(kind="uri", uri="document.pdf"), config=config)result = output.results[0]print(result.content)import { ExtractInputKind, extract } from '@xberg-io/xberg';
const output = await extract( { kind: ExtractInputKind.Uri, uri: 'document.pdf' }, { forceOcr: true, ocr: { backend: 'candle-deepseek-ocr', language: ['en'], backendOptions: { device: 'auto' }, }, },);const result = output.results[0];console.log(result.content);use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig};use serde_json::json;
let config = ExtractionConfig { force_ocr: true, ocr: Some(OcrConfig { backend: "candle-deepseek-ocr".into(), language: vec!["en".to_string()], backend_options: Some(json!({"device": "auto"})), ..Default::default() }), ..Default::default()};let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?;let result = &output.results[0];println!("{}", result.content);xberg extract document.pdf --force-ocr true --ocr-backend candle-deepseek-ocr --ocr-backend-options '{"device":"auto"}'Supported languages: English, Chinese, Japanese, Korean, French, German, Spanish, Italian, Portuguese, Russian, Arabic, Hindi, Thai, Vietnamese, and others.
Model source: deepseek-ai/DeepSeek-OCR on the Hugging Face Hub, auto-downloaded at a checksum-pinned revision. Pass backend_options.model_id to use a different repository (requires an explicit hf_revision) or backend_options.model_path to point at pre-staged local weights.
Candle PaddleOCR-VL
Section titled “Candle PaddleOCR-VL”PaddleOCR-VL 1.6 — SigLIP vision encoder + Ernie-4.5 text decoder for lightweight multilingual document understanding. Markdown output.
Ships in the published package by default on native desktop/server platforms — no feature flag or custom build needed. Set ocr.backend = "candle-paddleocr-vl" and the model downloads automatically on first use (~2.5 GB), cached at ~/.cache/huggingface/. Not available on WebAssembly, Android, iOS, Dart, or Swift.
For a custom Rust crate build, the model is gated behind the candle-paddleocr-vl feature (or the candle-vlm-ocr umbrella feature):
[dependencies]xberg = { version = "1", features = ["candle-paddleocr-vl"] }GPU support:
- Metal (macOS) — Default, F32 dtype
- CUDA (Linux/Windows with NVIDIA GPU) — Auto-detected
- CPU fallback — Slowest, but always available
Using Candle PaddleOCR-VL
Section titled “Using Candle PaddleOCR-VL”from xberg import ExtractInput, ExtractionConfig, OcrConfig, extract
config = ExtractionConfig( force_ocr=True, ocr=OcrConfig( backend="candle-paddleocr-vl", language=["en"], backend_options={"device": "auto", "model_path": "~/.cache/huggingface/"}, ),)output = await extract(ExtractInput(kind="uri", uri="document.pdf"), config=config)result = output.results[0]print(result.content)import { ExtractInputKind, extract } from '@xberg-io/xberg';
const output = await extract( { kind: ExtractInputKind.Uri, uri: 'document.pdf' }, { forceOcr: true, ocr: { backend: 'candle-paddleocr-vl', language: ['en'], backendOptions: { device: 'auto', model_path: '~/.cache/huggingface/' }, }, },);const result = output.results[0];console.log(result.content);use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig};use serde_json::json;
let config = ExtractionConfig { force_ocr: true, ocr: Some(OcrConfig { backend: "candle-paddleocr-vl".into(), language: vec!["en".to_string()], backend_options: Some(json!({"device": "auto", "model_path": "~/.cache/huggingface/"})), ..Default::default() }), ..Default::default()};let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?;let result = &output.results[0];println!("{}", result.content);xberg extract document.pdf --force-ocr true --ocr-backend candle-paddleocr-vl --ocr-backend-options '{"device":"auto","model_path":"~/.cache/huggingface/"}'Supported languages: English, Chinese, Japanese, Korean, French, German, Spanish, Italian, Portuguese, Russian, and others.
Model source: Download from PaddlePaddle Hub.
Using VLM OCR
Section titled “Using VLM OCR”Use a vision-language model (e.g. GPT-4o, Claude) as the OCR backend — each page is rendered and sent to the VLM. Cloud providers need an API key; local engines (Ollama, etc.) use the ollama/ prefix — see Local LLM Support.
import asynciofrom xberg import ExtractInput, extract, ExtractionConfig, OcrConfig, LlmConfig
async def main() -> None: config = ExtractionConfig( force_ocr=True, ocr=OcrConfig( backend="vlm", vlm_config=LlmConfig(model="openai/gpt-4o-mini"), ), ) result = await extract(ExtractInput(uri="scan.pdf"), config) print(result.results[0].content)
asyncio.run(main())import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = { forceOcr: true, ocr: { backend: "vlm", vlmConfig: { model: "openai/gpt-4o-mini", }, },};
const output = await extract({ kind: ExtractInputKind.Uri, uri: "scan.pdf" }, config);console.log(output.results?.[0]?.content);use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig, LlmConfig};
let config = ExtractionConfig { force_ocr: true, ocr: Some(OcrConfig { backend: "vlm".to_string(), vlm_config: Some(LlmConfig { model: "openai/gpt-4o-mini".to_string(), ..Default::default() }), ..Default::default() }), ..Default::default()};let output = extract(ExtractInput::from_uri("scan.pdf"), &config).await?;let result = &output.results[0];xberg extract scan.pdf --force-ocr true --vlm-model openai/gpt-4o-miniforce_ocr = true
[ocr]backend = "vlm"
[ocr.vlm_config]model = "openai/gpt-4o-mini"For more on VLM OCR, including custom prompts, supported providers, and API key configuration, see LLM Integration.
DPI Configuration
Section titled “DPI Configuration”Higher DPI improves accuracy but increases processing time and memory.
| DPI | Trade-off |
|---|---|
| 150 | Fastest — lower accuracy, less memory |
| 300 (default) | Balanced — good accuracy, reasonable speed |
| 600 | Best accuracy — slower, more memory |
import asynciofrom xberg import ( ExtractInput, extract, ExtractionConfig, OcrConfig, TesseractConfig, ImagePreprocessingConfig,)
async def main() -> None: config: ExtractionConfig = ExtractionConfig( ocr=OcrConfig( backend="tesseract", tesseract_config=TesseractConfig( preprocessing=ImagePreprocessingConfig(target_dpi=300), ), ), )
result = await extract(ExtractInput(uri="scanned.pdf"), config)
content_length: int = len(result.results[0].content) table_count: int = len(result.results[0].tables)
print(f"Content length: {content_length} characters") print(f"Tables detected: {table_count}")
asyncio.run(main())import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = { ocr: { backend: "tesseract", }, pdfOptions: { extractImages: true, },};
const output = await extract({ kind: ExtractInputKind.Uri, uri: "scanned.pdf" }, config);console.log(output.results?.[0]?.content);use xberg::{extract, ExtractionConfig, ExtractInput, ImagePreprocessingConfig, OcrConfig, TesseractConfig};
#[tokio::main]async fn main() -> xberg::Result<()> { let config = ExtractionConfig { ocr: Some(OcrConfig { backend: "tesseract".to_string(), tesseract_config: Some(TesseractConfig { preprocessing: Some(ImagePreprocessingConfig { target_dpi: 300, ..Default::default() }), ..Default::default() }), ..Default::default() }), ..Default::default() };
let _output = extract(ExtractInput::from_uri("scanned.pdf"), &config).await?; Ok(())}package main
import ( "log"
"github.com/xberg-io/xberg/packages/go")
func main() { targetDpi := int32(300) config := xberg.ExtractionConfig{ Ocr: &xberg.OcrConfig{ Backend: xberg.Ptr("tesseract"), TesseractConfig: &xberg.TesseractConfig{ Preprocessing: &xberg.ImagePreprocessingConfig{ TargetDpi: &targetDpi, }, }, }, } input := xberg.ExtractInputFromURI("scanned.pdf") result, err := xberg.Extract(*input, config) if err != nil { log.Fatalf("extract failed: %v", err) }
log.Println("content length:", len(result.Results[0].Content))}import io.xberg.Xberg;import io.xberg.ExtractInputKind;import io.xberg.ExtractionResult;import io.xberg.ExtractedDocument;import io.xberg.ExtractionConfig;import io.xberg.ExtractInput;import io.xberg.OcrConfig;import io.xberg.TesseractConfig;import io.xberg.ImagePreprocessingConfig;
ExtractionConfig config = ExtractionConfig.builder() .withOcr(OcrConfig.builder() .withBackend("tesseract") .withTesseractConfig(TesseractConfig.builder() .withPreprocessing(ImagePreprocessingConfig.builder() .withTargetDpi(300) .build()) .build()) .build()) .build();ExtractionResult output = Xberg.extract( ExtractInput.builder().withKind(ExtractInputKind.URI).withUri("scanned.pdf").build(), config);ExtractedDocument result = output.results().get(0);require 'xberg'
config = Xberg::ExtractionConfig.new( ocr: Xberg::OcrConfig.new(backend: 'tesseract'), pdf: Xberg::PdfConfig.new(dpi: 300))
input = Xberg::ExtractInput.new(uri: 'scanned.pdf')result = Xberg.extract(input, config)Advanced OCR Configuration
Section titled “Advanced OCR Configuration”Beyond backend and DPI selection, OcrConfig and ExtractionConfig expose finer control over when OCR runs, page orientation correction, the native-text-to-OCR fallback decision, multi-backend fallback, and structured element output. See the OcrConfig reference for every field.
Force OCR on Specific Pages
Section titled “Force OCR on Specific Pages”Set force_ocr_pages on ExtractionConfig to OCR only the listed pages (1-indexed). Unlisted pages use native text extraction. PDF only. Ignored when force_ocr is true; duplicates are deduplicated. Provide an ocr config for backend and language selection — defaults are used if absent.
from xberg import ExtractInput, ExtractionConfig, extract
config = ExtractionConfig(force_ocr_pages=[1, 3, 5])output = await extract(ExtractInput(kind="uri", uri="document.pdf"), config=config)result = output.results[0]import { ExtractInputKind, extract } from '@xberg-io/xberg';
const output = await extract( { kind: ExtractInputKind.Uri, uri: 'document.pdf' }, { forceOcrPages: [1, 3, 5] },);const result = output.results[0];use xberg::{extract, ExtractInput, ExtractionConfig};
let config = ExtractionConfig { force_ocr_pages: Some(vec![1, 3, 5]), ..Default::default()};let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?;let result = &output.results[0];force_ocr_pages = [1, 3, 5]Auto-Rotate Pages
Section titled “Auto-Rotate Pages”Set auto_rotate on OcrConfig to detect page orientation (0/90/180/270 degrees) with Tesseract’s orientation-and-script detection before recognition. Pages rotated with high confidence are corrected before OCR — important for rotated scans. Defaults to false.
from xberg import ExtractInput, ExtractionConfig, OcrConfig, extract
config = ExtractionConfig( force_ocr=True, ocr=OcrConfig(backend="tesseract", auto_rotate=True),)output = await extract(ExtractInput(kind="uri", uri="rotated_scan.pdf"), config=config)result = output.results[0]import { ExtractInputKind, extract } from '@xberg-io/xberg';
const output = await extract( { kind: ExtractInputKind.Uri, uri: 'rotated_scan.pdf' }, { forceOcr: true, ocr: { backend: 'tesseract', autoRotate: true } },);const result = output.results[0];use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig};
let config = ExtractionConfig { force_ocr: true, ocr: Some(OcrConfig { backend: "tesseract".into(), auto_rotate: true, ..Default::default() }), ..Default::default()};let output = extract(ExtractInput::from_uri("rotated_scan.pdf"), &config).await?;let result = &output.results[0];[ocr]backend = "tesseract"auto_rotate = trueQuality Thresholds
Section titled “Quality Thresholds”The native-text-to-OCR fallback decision (should a PDF page with an embedded text layer be re-OCR’d?) and multi-backend stage acceptance are governed by quality_thresholds (OcrQualityThresholds) on OcrConfig. When unset, compiled defaults apply. Frequently tuned fields:
| Field | Default | Purpose |
|---|---|---|
pipeline_min_quality |
0.5 |
Minimum quality score (0.0-1.0) for a pipeline stage result to be accepted; below this, the next backend runs. |
min_meaningful_words |
3 |
Minimum count of meaningful words before native text is accepted instead of OCR. |
min_alnum_ratio |
0.3 |
Minimum alphanumeric ratio of non-whitespace characters. |
critical_fragmented_word_ratio |
0.80 |
Fraction of 1-2 character words that forces OCR regardless of other signals. |
max_ocr_output_fragmented_word_ratio |
0.35 |
Fraction of short words that reports OCR output as suspected recognition noise. |
min_ocr_mean_confidence |
75.0 |
Calibrated backend confidence below which OCR output is reported as suspected noise. |
discard_suspected_ocr_noise |
false |
Discard non-empty OCR output when a recognition-noise signal fires. By default, Xberg retains the text and adds a processing warning. |
min_undecodable_ratio |
0.5 |
Minimum fraction of non-whitespace characters that must be undecodable (PUA, replacement, or control garbage) before a page’s text layer is treated as unreadable and routed to OCR. See Automatic OCR for Undecodable Text Layers. |
min_reliable_language_chunk_ratio |
0.10 |
Minimum fraction of a page’s prose chunks whatlang must classify as a reliable language before the page is trusted as legible. See Automatic OCR for Language/Dictionary-Implausible Text. |
Inspect processing_warnings for pages flagged by fragmented-word, calibrated-confidence, or configured
dictionary-invalid-word signals. To preserve the pre-1.1 destructive behavior, set
discard_suspected_ocr_noise = true; blank OCR pages remain blank without a recognition-noise warning.
from xberg import ExtractionConfig, OcrConfig, OcrQualityThresholds, extract
config = ExtractionConfig( ocr=OcrConfig( quality_thresholds=OcrQualityThresholds(pipeline_min_quality=0.6), ),)use xberg::{ExtractionConfig, OcrConfig, OcrQualityThresholds};
let config = ExtractionConfig { ocr: Some(OcrConfig { quality_thresholds: Some(OcrQualityThresholds { pipeline_min_quality: 0.6, ..Default::default() }), ..Default::default() }), ..Default::default()};[ocr.quality_thresholds]pipeline_min_quality = 0.6# Opt in only when discarding suspected OCR noise is preferable to retaining marginal text.discard_suspected_ocr_noise = trueAutomatic OCR for Undecodable Text Layers
Section titled “Automatic OCR for Undecodable Text Layers”PDFs with a font whose CID/glyph indices have no ToUnicode mapping (common with subset Identity-H fonts) often “extract” successfully but decode to garbage — Unicode Private Use Area (PUA) codepoints, replacement characters (U+FFFD), or non-whitespace control characters instead of legible text. Previously this garbage passed the native-text quality check and OCR never ran. Xberg now detects this case and routes the page to OCR automatically, the same as a scanned page (issue #1254).
A page is treated as having an undecodable text layer, and routed to OCR under OcrStrategy.Auto, when:
- it has at least
min_total_non_whitespace(default64) non-whitespace characters — short snippets with a stray symbol don’t trip this check, and - at least
min_undecodable_ratio(default0.5) of those non-whitespace characters are undecodable (PUA, U+FFFD, or non-whitespace control characters).
Tune the ratio via OcrQualityThresholds.min_undecodable_ratio on OcrConfig.quality_thresholds:
from xberg import ExtractInput, ExtractionConfig, OcrConfig, OcrQualityThresholds, extract
config = ExtractionConfig( ocr=OcrConfig( quality_thresholds=OcrQualityThresholds(min_undecodable_ratio=0.3), ),)output = await extract(ExtractInput(kind="uri", uri="identity-h-subset-font.pdf"), config=config)result = output.results[0]# Pages whose text layer decoded to >=30% PUA/replacement/control garbage were OCR'd.use xberg::{ExtractionConfig, OcrConfig, OcrQualityThresholds};
let config = ExtractionConfig { ocr: Some(OcrConfig { quality_thresholds: Some(OcrQualityThresholds { min_undecodable_ratio: 0.3, ..Default::default() }), ..Default::default() }), ..Default::default()};[ocr.quality_thresholds]min_undecodable_ratio = 0.3Automatic OCR for Language/Dictionary-Implausible Text
Section titled “Automatic OCR for Language/Dictionary-Implausible Text”A /ToUnicode CMap (or other mapping tier) can resolve every glyph to a character, but consistently the WRONG one — a ROT-shifted mapping is the clearest example. The decoded text is structurally indistinguishable from real prose: normal alphanumeric ratio, normal word lengths, no fragmentation, and no undecodable (PUA/replacement/control) characters, so neither the character-shape checks above nor the undecodable-text-layer check can see the problem. Xberg detects this by checking whether the decoded text reads as any real, detectable language at all (issue #1696).
A page is treated as language/dictionary-implausible, and routed to OCR under OcrStrategy.Auto, when:
- it has enough prose content to judge in the first place — pages that are mostly numeric tables, formulas, or code abstain rather than guess;
- less than
min_reliable_language_chunk_ratio(default0.10) of its prose chunks are classified as a reliable language; and - the mean detection confidence across those chunks is also low, so a genuinely legible page in a close language pair (e.g. Danish vs. Norwegian Bokmål) is not misclassified.
A page the character-shape or undecodable-text-layer checks already route to OCR is not independently re-flagged by this signal. PdfMetadata.implausible_text_pages reports the 1-indexed pages this check flagged; it is None when enable_plausibility_ocr_routing is false. If no page holds enough prose to judge, the check reaches no verdict and the list stays empty for that reason, not because the text layer is sound. Xberg then records a processing warning naming how many pages it examined, so an empty list is not mistaken for a clean result (issue #1709).
Tune the ratio via OcrQualityThresholds.min_reliable_language_chunk_ratio on OcrConfig.quality_thresholds, or disable the check with enable_plausibility_ocr_routing:
from xberg import ExtractInput, ExtractionConfig, OcrConfig, OcrQualityThresholds, extract
config = ExtractionConfig( ocr=OcrConfig( quality_thresholds=OcrQualityThresholds(min_reliable_language_chunk_ratio=0.2), ),)output = await extract(ExtractInput(kind="uri", uri="wrong-mapping.pdf"), config=config)result = output.results[0]# Pages whose decoded text did not read as any real language were OCR'd.use xberg::{ExtractionConfig, OcrConfig, OcrQualityThresholds};
let config = ExtractionConfig { ocr: Some(OcrConfig { quality_thresholds: Some(OcrQualityThresholds { min_reliable_language_chunk_ratio: 0.2, ..Default::default() }), ..Default::default() }), ..Default::default()};[ocr.quality_thresholds]min_reliable_language_chunk_ratio = 0.2Multi-Backend Pipeline
Section titled “Multi-Backend Pipeline”Set pipeline (OcrPipelineConfig) to try several backends in priority order (highest first) with quality-based fallback. Each stage’s output is scored; if it meets pipeline_min_quality, it is accepted, otherwise the next stage runs. Each stage (OcrPipelineStage) has a backend and priority (default 100), plus optional language, tesseract_config, paddle_ocr_config, vlm_config, and backend_options.
When pipeline is set, the top-level backend, vlm_fallback, and backend_options are ignored — configure each stage directly instead.
from xberg import OcrConfig, OcrPipelineConfig, OcrPipelineStage, OcrQualityThresholds
ocr = OcrConfig( pipeline=OcrPipelineConfig( stages=[ OcrPipelineStage(backend="tesseract", priority=100), OcrPipelineStage(backend="paddleocr", priority=50), ], quality_thresholds=OcrQualityThresholds(pipeline_min_quality=0.6), ),)use xberg::{OcrConfig, OcrPipelineConfig, OcrPipelineStage, OcrQualityThresholds};
let ocr = OcrConfig { pipeline: Some(OcrPipelineConfig { stages: vec![ OcrPipelineStage { backend: "tesseract".into(), priority: 100, language: None, tesseract_config: None, paddle_ocr_config: None, vlm_config: None, backend_options: None, }, OcrPipelineStage { backend: "paddleocr".into(), priority: 50, language: None, tesseract_config: None, paddle_ocr_config: None, vlm_config: None, backend_options: None, }, ], quality_thresholds: OcrQualityThresholds { pipeline_min_quality: 0.6, ..Default::default() }, }), ..Default::default()};[[ocr.pipeline.stages]]backend = "tesseract"priority = 100
[[ocr.pipeline.stages]]backend = "paddleocr"priority = 50
[ocr.pipeline.quality_thresholds]pipeline_min_quality = 0.6VLM Fallback
Section titled “VLM Fallback”vlm_fallback (VlmFallbackPolicy) is ergonomic sugar over an explicit pipeline. When set and pipeline is None, an equivalent pipeline is synthesised automatically. It requires vlm_config to be set. When pipeline is explicitly set, vlm_fallback is ignored.
disabled(default) — single-backend mode.on_low_quality— run the classical backend first; if the result scores belowquality_threshold, retry the page with the VLM.always— skip the classical backend; send every page to the VLM.
from xberg import LlmConfig, OcrConfig, VlmFallbackPolicy
ocr = OcrConfig( vlm_fallback=VlmFallbackPolicy.on_low_quality(0.6), vlm_config=LlmConfig(model="openai/gpt-4o-mini"),)use xberg::{LlmConfig, OcrConfig, VlmFallbackPolicy};
let ocr = OcrConfig { vlm_fallback: VlmFallbackPolicy::OnLowQuality { quality_threshold: 0.6 }, vlm_config: Some(LlmConfig { model: "openai/gpt-4o-mini".into(), ..Default::default() }), ..Default::default()};[ocr.vlm_fallback]mode = "on_low_quality"quality_threshold = 0.6
[ocr.vlm_config]model = "openai/gpt-4o-mini"OCR Elements and Word-Level Bounding Boxes
Section titled “OCR Elements and Word-Level Bounding Boxes”Set element_config (OcrElementConfig) on OcrConfig to emit structured OCR elements with spatial and confidence data. When include_elements is true, the result’s ocr_elements field is populated with OcrElement entries — each carries text, geometry (a rectangle with left/top/width/height in pixels, or a 4-point quadrilateral for rotated text), confidence (detection and recognition, 0.0-1.0), level, optional rotation, and a 1-indexed page_number.
include_elements— populateocr_elements. Defaults tofalse.min_level— minimum hierarchy level to include:word,line(default),block, orpage. Elements below this level are dropped.min_confidence— drop elements with recognition confidence below this (0.0-1.0).build_hierarchy— populateparent_idfrom spatial containment (Tesseract only).
from xberg import ( ExtractInput, ExtractionConfig, OcrConfig, OcrElementConfig, OcrElementLevel, extract,)
config = ExtractionConfig( force_ocr=True, ocr=OcrConfig( element_config=OcrElementConfig( include_elements=True, min_level=OcrElementLevel.WORD, min_confidence=0.5, ), ),)output = await extract(ExtractInput(kind="uri", uri="scan.png"), config=config)result = output.results[0]for element in result.ocr_elements or []: print(element.text, element.confidence.recognition, element.page_number)import { ExtractInputKind, OcrElementLevel, extract } from '@xberg-io/xberg';
const output = await extract( { kind: ExtractInputKind.Uri, uri: 'scan.png' }, { forceOcr: true, ocr: { elementConfig: { includeElements: true, minLevel: OcrElementLevel.Word, minConfidence: 0.5, }, }, },);const result = output.results[0];for (const element of result.ocrElements ?? []) { console.log(element.text, element.confidence?.recognition, element.pageNumber);}use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig, OcrElementConfig, OcrElementLevel};
let config = ExtractionConfig { force_ocr: true, ocr: Some(OcrConfig { element_config: Some(OcrElementConfig { include_elements: true, min_level: OcrElementLevel::Word, min_confidence: 0.5, build_hierarchy: false, }), ..Default::default() }), ..Default::default()};let output = extract(ExtractInput::from_uri("scan.png"), &config).await?;let result = &output.results[0];for element in result.ocr_elements.iter().flatten() { println!("{} {} p{}", element.text, element.confidence.recognition, element.page_number);}PP-OCRv6
Section titled “PP-OCRv6”PP-OCRv6 is the default PaddleOCR model generation (model_version defaults to "pp-ocrv6"). It adds a unified CJK+Latin+Japanese/Korean recognition model with three size tiers, selected via model_tier:
| Tier | Notes |
|---|---|
medium |
~62MB detection, full 18,708-char CJK+Latin+JA/KO dictionary. Highest accuracy, substantially slower on CPU. |
small |
Default. ~9.9MB detection, same dictionary as medium. |
tiny |
~1.8MB detection, but a reduced 6,904-char (~zh/en) dictionary — it cannot read the scripts the other tiers cover. |
model_tier defaults to "mobile", which auto-resolves to the v6 "small" tier — you don’t need to set model_tier explicitly to get v6’s default behavior. A legacy "server" tier value, or any unrecognised value, resolves to "medium".
PaddleOCR pages do not run concurrently — the ONNX session is held behind a mutex, so the thread budget goes to intra-op parallelism and wall time scales with page count times per-page inference. On a long document the tier choice, not the thread count, is what governs how long extraction takes.
Scripts outside v6’s unified coverage — Arabic, Cyrillic, Devanagari, Greek, Tamil, Telugu, and Thai — transparently fall back to the PP-OCRv5 per-script recognition models; no configuration change is needed to use those languages under v6.
To pin the legacy PP-OCRv5 fleet (per-script + unified recognition models, mobile/server tiers), set model_version to "pp-ocrv5" explicitly:
from xberg import ExtractInput, ExtractionConfig, OcrConfig, extract
config = ExtractionConfig( ocr=OcrConfig( backend="paddleocr", backend_options={"model_version": "pp-ocrv5", "model_tier": "mobile"}, ),)output = await extract(ExtractInput(kind="uri", uri="scanned.pdf"), config=config)use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig};use serde_json::json;
let config = ExtractionConfig { ocr: Some(OcrConfig { backend: "paddleocr".into(), backend_options: Some(json!({"model_version": "pp-ocrv5", "model_tier": "mobile"})), ..Default::default() }), ..Default::default()};let output = extract(ExtractInput::from_uri("scanned.pdf"), &config).await?;PaddleOCR-VL 1.6 (candle backend)
Section titled “PaddleOCR-VL 1.6 (candle backend)”The candle-paddleocr-vl backend now defaults to PaddleOCR-VL 1.6 (SigLIP vision encoder + Ernie-4.5 text decoder), auto-downloading xberg-io/paddleocr-vl-1.6 — a checksum-pinned mirror of PaddlePaddle/PaddleOCR-VL-1.6 — on first use. It runs one of four tasks via backend_options.task: ocr (default), table, formula, or chart. See Candle PaddleOCR-VL above for installation and usage.
PaddleOCR Script Families
Section titled “PaddleOCR Script Families”80+ languages across 11 script families (PP-OCRv5). Recognition models are downloaded on demand from HuggingFace:
| Family | Languages |
|---|---|
| English | English, numbers, punctuation |
| Chinese | Simplified/Traditional Chinese, Japanese |
| Latin | French, German, Spanish, Portuguese, Italian, Polish, Dutch, Turkish, Vietnamese, and so on. |
| Korean | Korean (Hangul) |
| Slavic | Russian, Ukrainian, Belarusian, Bulgarian, Serbian, and so on. |
| Thai | Thai script |
| Greek | Greek script |
| Arabic | Arabic, Persian, Urdu |
| Devanagari | Hindi, Marathi, Sanskrit, Nepali |
| Tamil | Tamil script |
| Telugu | Telugu script |
Models are cached locally after first download, so subsequent runs start immediately.
CLI Usage
Section titled “CLI Usage”# Basic OCR extractionxberg extract scanned.pdf --ocr true
# Specific languagexberg extract french_doc.pdf --ocr true --ocr-language fra
# Specific backendxberg extract chinese_doc.pdf --ocr true --ocr-backend paddle-ocr --ocr-language ch
# Sceptre on a native desktop or serverxberg extract german_doc.pdf --ocr true --ocr-backend sceptre --ocr-language deu
# Force OCR on all pagesxberg extract document.pdf --force-ocr true
# VLM OCR backendxberg extract handwritten.pdf --force-ocr true --vlm-model openai/gpt-4o-mini
# Use a config filexberg extract scanned.pdf --config xberg.toml --ocr true| Flag | Description |
|---|---|
--ocr true |
Enable OCR processing |
--ocr-language <code> |
Language code (eng, deu, fra, ch, ja, ru, etc.) |
--ocr-backend <backend> |
Engine: tesseract, paddle-ocr, sceptre, a candle-* backend, or vlm |
--force-ocr true |
OCR all pages regardless of text layer |
--vlm-model <model> |
VLM model for OCR (for example, openai/gpt-4o-mini). Implies --ocr-backend vlm |
Troubleshooting
Section titled “Troubleshooting”Next Steps
Section titled “Next Steps”- LLM Integration — VLM OCR, structured extraction, and LLM embeddings
- Configuration — all configuration options
- Extraction Basics — core extraction API and supported formats
- Chunking — split text for RAG
- Language Detection — multilingual document analysis
- Embeddings — semantic vectors for search