Skip to content

Document Summarisation

Generate a one-paragraph summary of extracted documents for search snippets, indexing, or quick reviews. Choose extractive summarization for deterministic, network-free local processing, or abstractive for fluent, AI-generated prose.

Strategy Cargo feature Network Quality Latency
Extractive (default) summarization None — fully local Sentence-level selection from source < 100 ms typical
Abstractive summarization-llm LLM provider Generates novel prose, can summarise across sentences Provider-dependent
  • You need a one-paragraph TL;DR for indexing or search snippets.
  • You need a deterministic, network-free summary (extractive only).
  • You need a fluent abstractive summary for downstream LLM consumption.
  • You need full per-section summaries. Chunk the document first and summarise each chunk separately.
  • You need cross-document summarisation. Summarise per document, then summarise the summaries with the LLM backend.

TextRank extractive summary over a multi-paragraph plain text document. Pure-Rust, deterministic, no external services required.

Python
import asyncio
from xberg import extract, ExtractInput, ExtractInputKind
from xberg._xberg import ExtractionConfig
async def main() -> None:
input = ExtractInput(kind=ExtractInputKind("uri"), uri="https://example.com/text/book_war_and_peace_1p.txt")
config = ExtractionConfig.from_json("{\"summarization\":{\"max_tokens\":80,\"strategy\":\"extractive\"}}")
result = await extract(input, config)
print(result.results[0].summary)
asyncio.run(main())

For CLI and server configuration, add the same settings to xberg.toml:

xberg.toml
[summarization]
strategy = "extractive"
max_tokens = 200

Switch the strategy and attach an LlmConfig:

LLM-driven abstractive summary. Skipped automatically when XBERG_LLM_API_KEY (or OPENAI_API_KEY) is not set.

Python
import asyncio
from xberg import extract, ExtractInput, ExtractInputKind
from xberg._xberg import ExtractionConfig
async def main() -> None:
input = ExtractInput(kind=ExtractInputKind("uri"), uri="https://example.com/text/book_war_and_peace_1p.txt")
config = ExtractionConfig.from_json("{\"summarization\":{\"llm\":{\"max_tokens\":200,\"model\":\"openai/gpt-4o-mini\",\"temperature\":0.0},\"max_tokens\":150,\"strategy\":\"abstractive\"}}")
result = await extract(input, config)
print(result.results[0].summary)
asyncio.run(main())

The model receives the extracted content and returns the summary verbatim. Token usage records in ExtractedDocument.llm_usage with source = "summarisation_abstractive".

Strategy What max_tokens caps
Extractive Loose whitespace tokens in the output summary. The TextRank selector stops appending sentences once it would exceed the cap.
Abstractive A prompt hint asking the model for approximately this many tokens — not a provider hard cap. The provider’s request limit comes separately from SummarizationConfig.llm.max_tokens.

Leave None to let the backend pick a sensible default.

{
"summary": {
"text": "The contract sets out a 3-year support agreement with quarterly billing and a fixed escalation cap of 4%.",
"strategy": "extractive",
"token_count": 19
}
}

Pick any liter-llm provider — see LLM Integration. For most documents, gpt-4o-mini, claude-3-5-haiku, or google/gemini-2.0-flash give good cost / quality trade-offs.

API-key precedence:

  1. SummarizationConfig.llm.api_key
  2. XBERG_LLM_API_KEY
  3. Per-provider env var