Skip to content

LLM Integration

Xberg integrates with 165 LLM providers (including local inference engines) via liter-llm for three capabilities: VLM OCR, structured extraction, and provider-hosted embeddings.

Use vision-language models as an OCR backend by rendering document pages as images and sending them to the VLM for text extraction.

  • Low-quality scanned documents where traditional OCR struggles
  • Handwritten text recognition
  • Arabic, Farsi, and other scripts with poor Tesseract/PaddleOCR support
  • Complex layouts where traditional OCR fails (mixed tables, forms, diagrams)
  • When you need higher accuracy and can accept higher latency and API costs
Python
import asyncio
from xberg import ExtractInput, extract, ExtractionConfig, OcrConfig, LlmConfig
async def main() -> None:
config = ExtractionConfig(
force_ocr=True,
ocr=OcrConfig(
backend="vlm",
vlm_config=LlmConfig(model="openai/gpt-4o-mini"),
),
)
result = await extract(ExtractInput(uri="scan.pdf"), config)
print(result.results[0].content)
asyncio.run(main())

Override the default prompt template for VLM OCR:

Python
from xberg import ExtractionConfig, OcrConfig, LlmConfig
config = ExtractionConfig(
force_ocr=True,
ocr=OcrConfig(
backend="vlm",
vlm_config=LlmConfig(model="openai/gpt-4o-mini"),
vlm_prompt="Extract all text from this document image. Preserve formatting.",
),
)

Any liter-llm vision-capable provider works as a VLM OCR backend:

Provider Example Model
OpenAI openai/gpt-4o, openai/gpt-4o-mini
Anthropic anthropic/claude-3-5-sonnet-20241022
Google google/gemini-2.0-flash
Groq groq/llama-3.2-90b-vision-preview
Ollama (local) ollama/llama3.2-vision
LM Studio (local) lmstudio/llava-1.5
vLLM (local) vllm/llava-next

Extract structured JSON data from documents by providing a schema; the document text is sent to an LLM for conforming extraction.

Python
import asyncio
import json
from xberg import ExtractInput, extract, ExtractionConfig, StructuredExtractionConfig, LlmConfig
async def main() -> None:
config = ExtractionConfig(
structured_extraction=StructuredExtractionConfig(
schema=json.dumps({
"type": "object",
"properties": {
"title": {"type": "string"},
"authors": {"type": "array", "items": {"type": "string"}},
"date": {"type": "string"},
},
"required": ["title", "authors", "date"],
"additionalProperties": False,
}),
schema_name="paper",
llm=LlmConfig(model="openai/gpt-4o-mini"),
strict=True,
),
)
result = await extract(ExtractInput(uri="paper.pdf"), config)
print(result.results[0].structured_output)
# {"title": "...", "authors": ["..."], "date": "..."}
asyncio.run(main())

Override the default extraction prompt with a Jinja2 template:

Python
from xberg import ExtractionConfig, StructuredExtractionConfig, LlmConfig
config = ExtractionConfig(
structured_extraction=StructuredExtractionConfig(
schema={"type": "object", "properties": {"title": {"type": "string"}}},
llm=LlmConfig(model="openai/gpt-4o-mini"),
prompt=(
"Analyze this document and extract key metadata.\n\n"
"Document:\n{{ content }}\n\n"
"Schema: {{ schema }}"
),
),
)

Available template variables:

Variable Description
{{ content }} The extracted document text
{{ schema }} The JSON schema as a formatted string
{{ schema_name }} The schema name (default: "extraction")
{{ schema_description }} The schema description (may be empty)

Structured extraction handles provider differences automatically:

  • OpenAI: Full strict mode with additionalProperties enforcement
  • Anthropic/Gemini: additionalProperties automatically stripped (not supported by these providers)
  • All providers: Markdown code fence wrapping in responses is automatically handled

When strict=True, the LLM is instructed to produce output that exactly matches the schema. This enables OpenAI’s structured output mode and adds validation on the response.

Use provider-hosted embedding models when you need to match your vector database model or local ONNX models are unavailable.

Python
import asyncio
from xberg import embed, EmbeddingConfig, EmbeddingModelType, LlmConfig
async def main() -> None:
config = EmbeddingConfig(
model=EmbeddingModelType.llm(
LlmConfig(model="openai/text-embedding-3-small")
),
normalize=True,
)
embeddings = await embed(["Hello world"], config=config)
print(len(embeddings[0])) # 1536
asyncio.run(main())
Model Dimensions Provider
openai/text-embedding-3-small 1536 OpenAI
openai/text-embedding-3-large 3072 OpenAI
mistral/mistral-embed 1024 Mistral
Any liter-llm embedding-capable provider Varies Various

Run local LLM inference engines via liter-llm’s provider routing; point to your local server without needing an API key.

Engine Prefix Default URL Install
Ollama ollama/ http://localhost:11434/v1 brew install ollama
LM Studio lmstudio/ http://localhost:1234/v1 Desktop app
vLLM vllm/ http://localhost:8000/v1 pip install vllm
llama.cpp llamacpp/ http://localhost:8080/v1 Build from source
LocalAI localai/ http://localhost:8080/v1 Docker
llamafile llamafile/ http://localhost:8080/v1 Single binary
Terminal window
# Start Ollama and pull a model
ollama pull llama3.2-vision
# Use it for VLM OCR (no API key needed)
xberg extract scan.pdf --force-ocr true \
--vlm-model ollama/llama3.2-vision
# Use it for structured extraction
xberg extract doc.pdf --config structured-extraction.toml --format json
# Use it for embeddings
xberg embed --provider llm \
--model ollama/all-minilm \
--text "Hello world"

Every LLM call made during extraction is tracked in the llm_usage field of ExtractedDocument. Each entry records the model used, token counts, estimated cost, and why the model stopped generating.

from xberg import ExtractInput, extract
output = await extract(ExtractInput(kind="uri", uri="document.pdf"), config)
result = output.results[0]
if result.get("llm_usage"):
for usage in result["llm_usage"]:
print(f"{usage['source']}: {usage['input_tokens']} in, {usage['output_tokens']} out, ${usage['estimated_cost']:.4f}")

The source field indicates which pipeline stage triggered the call: "vlm_ocr", "structured_extraction", or "embeddings".

This guide is the canonical reference for LLM API-key precedence.

When Xberg builds an LLM client, the key is resolved in this order (highest priority first):

  1. api_key field on the feature’s LlmConfig (VLM OCR vlm_config, structured_extraction.llm, or the embedding LlmConfig). If set, it is used verbatim.
  2. Provider standard env var (OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY, etc.), resolved by liter-llm when api_key is unset.

The XBERG_LLM_API_KEY env var is not a general per-provider fallback. It is read only during CLI/server config loading and populates structured_extraction.llm.api_key. Because it fills the api_key field, it overrides a config-file value and takes precedence over the provider standard env var for structured extraction. The same env var (along with XBERG_LLM_BASE_URL) is also forwarded onto ocr.vlm_config.api_key / ocr.vlm_config.base_url when a VLM OCR backend is already configured (issue #1339) — it never enables the VLM path on its own. Embeddings have no Xberg-specific key env var; set api_key directly on the embedding LlmConfig or rely on the provider standard env var.

Python
from xberg import LlmConfig
# Explicit API key
config = LlmConfig(model="openai/gpt-4o", api_key="sk-...")
# Custom base URL (e.g., Azure OpenAI, local proxy)
config = LlmConfig(
model="openai/gpt-4o",
base_url="https://my-proxy.example.com/v1",
)
Field Type Default Description
model str required Provider/model in liter-llm format (for example, "openai/gpt-4o")
api_key str | None None API key (falls back to env vars)
base_url str | None None Custom endpoint URL
timeout_secs int | None 60 Request timeout in seconds (300s default for VLM OCR)
max_retries int | None 3 Maximum retry attempts
temperature float | None None Sampling temperature
max_tokens int | None None Maximum tokens to generate
load_env bool | None None Whether liter-llm loads provider credentials from environment vars
headers dict[str, str] | None None Extra HTTP headers sent with every request
providers list[LlmProviderConfig] None Custom OpenAI-compatible providers, routed by model prefix — see below
cache LlmCacheConfig | None None Response cache settings — requires liter-llm’s tower feature
budget LlmBudgetConfig | None None Spend limits and enforcement — requires liter-llm’s tower feature
rate_limit LlmRateLimitConfig | None None Requests/tokens per minute — requires liter-llm’s tower feature
cost_tracking bool | None None Per-request cost tracking — requires liter-llm’s tower feature
tracing bool | None None OpenTelemetry-compatible spans — requires liter-llm’s tower feature
cooldown_secs int | None None Cooldown after transient errors — requires liter-llm’s tower feature
health_check_secs int | None None Background health check interval — requires liter-llm’s tower feature
bedrock BedrockConfig | None None AWS region/credentials for bedrock/-prefixed models

Full field-by-field descriptions, including LlmProviderConfig, LlmCacheConfig, LlmBudgetConfig, LlmRateLimitConfig, and BedrockConfig, are in the Configuration Reference.

Fields That Require liter-llm’s tower Feature

Section titled “Fields That Require liter-llm’s tower Feature”

cache, budget, rate_limit, cost_tracking, tracing, cooldown_secs, and health_check_secs are passed straight through to liter-llm’s client::LlmConfig. They only take effect when liter-llm is compiled with its tower feature. Without it, Xberg accepts and round-trips the values (TOML load, JSON serialization, every language binding) but liter-llm ignores them — no error, no warning.

xberg.toml
[structured_extraction.llm]
model = "openai/gpt-4o"
cost_tracking = true
tracing = true
cooldown_secs = 30
health_check_secs = 60
[structured_extraction.llm.cache]
max_entries = 512
ttl_seconds = 600
backend = "memory"
[structured_extraction.llm.budget]
global_limit = 100.0
enforcement = "hard"
[structured_extraction.llm.budget.model_limits]
"openai/gpt-4o" = 25.0
[structured_extraction.llm.rate_limit]
rpm = 60
tpm = 100000
window_seconds = 60

LlmConfig.providers registers custom OpenAI-compatible endpoints and routes models to them by prefix. Every entry is registered with liter-llm when Xberg builds a client, so any model whose name starts with one of the declared model_prefixes is sent to that provider’s base_url.

xberg.toml
[structured_extraction.llm]
model = "my-provider/llama-3.1-70b"
api_key = "..."
[[structured_extraction.llm.providers]]
name = "my-provider"
base_url = "https://my-llm.example.com/v1"
auth_header = "X-Api-Key"
model_prefixes = ["my-provider/"]

auth_header is the header name carrying the key. Leave it unset (or set it to Authorization) for Authorization: Bearer <key>; any other name sends the raw key under that header, with no scheme prefix.

Registration failures surface as a validation error rather than being ignored. The liter-llm registry is process-global and keyed by name, so the most recently built client wins for a given provider name — use distinct names across every LlmConfig in one process.

For a single custom or self-hosted endpoint that needs no prefix routing, set base_url on LlmConfig directly instead (see Custom Base URL above).

Credentials Are Redacted in Debug, Not in Serialized Output

Section titled “Credentials Are Redacted in Debug, Not in Serialized Output”

api_key and the Bedrock credential fields (access_key_id, secret_access_key, session_token) never appear in Debug output or logs. They do appear in serialized TOML/JSON — that’s how a config file persists them across runs. This is by design, not a defect. If you write a config file containing these fields, handle it like any other secrets file: restrictive permissions, never world-readable, never committed.

Use the unified /extract endpoint or extract MCP tool with a structured_extraction config object. LLM-hosted embeddings are Rust-only for now and can be wired into extraction through EmbeddingConfig.