Xberg Document Extraction skill

Extract text, tables, metadata, and images from 107 document formats (PDF, Office, images, HTML, email, archives, academic) using Xberg.

by xberg-io·MIT license·★ 9,384 Stars on the repo·GitHub ↗

Use now

Files of Xberg Document Extraction

xberg-io/main1 file shown
SKILL.md
Show the full text428 lines

Xberg Document Extraction

Xberg is a document intelligence library with a Rust core and bindings for Python, TypeScript/Node.js, Ruby, PHP, Go, Java, C#, Elixir, WebAssembly, Dart, Kotlin Android, Swift, Zig, and C. It extracts text, tables, metadata, and images from 107 formats across 140 unique file extensions and accepts 53 compatibility MIME aliases, including PDF, Office documents, images, HTML, email, archives, and academic formats.

Use this skill when writing code that:

  • Extracts text or metadata from documents
  • Performs OCR on scanned documents or images
  • Batch-processes multiple files
  • Configures extraction options (output format, chunking, OCR, language detection)
  • Implements custom plugins (post-processors, validators, OCR backends)

If the xberg MCP server is registered in this session, prefer its tools over shelling out to the CLI — they expose the same extraction surface with structured arguments and results.

Installation

Python
pip install xberg
Node.js
npm install @xberg-io/xberg
Rust
cargo add xberg
# Cargo.toml
[dependencies]
xberg = { version = "1.1.0", features = ["full"] }
tokio = { version = "1", features = ["full"] }
# feature flags: pdf, ocr, chunking, embeddings, language-detection, keywords, api, mcp
#                (or "formats" / "full" aggregates); tokio-runtime is on by default
CLI
brew install xberg-io/tap/xberg
# or run without a persistent install (the CLI proxy package self-installs the binary):
npx @xberg-io/xberg-cli --help
uvx --from xberg-cli xberg --help
# or download a prebuilt binary from the latest GitHub release:
#   https://github.com/xberg-io/xberg/releases/latest
# or build from source:
cargo install xberg-cli

Quick Start

The library entry points are extract(input, config) and extract_batch(inputs, config). Both return an ExtractionResult envelope — the extracted document(s) live in result.results, and per-document data (content, tables, metadata, …) is on each result.results[i]. Python and Node are async-only.

Python
import asyncio
from xberg import ExtractInput, extract, ExtractionConfig

async def main() -> None:
    result = await extract(ExtractInput(uri="document.pdf"), ExtractionConfig())
    doc = result.results[0]
    print(doc.content)    # extracted text
    print(doc.metadata)   # document metadata
    print(doc.tables)     # extracted tables

asyncio.run(main())
Node.js
import { extract } from "@xberg-io/xberg";

const output = await extract({ kind: "uri", uri: "document.pdf" });
const doc = output.results[0];
console.log(doc.content);
console.log(doc.metadata);
console.log(doc.tables);
Rust
use xberg::{extract, ExtractInput, ExtractionConfig};

#[tokio::main]
async fn main() -> xberg::Result<()> {
    let output = extract(ExtractInput::from_uri("document.pdf"), &ExtractionConfig::default()).await?;
    println!("{}", output.results[0].content);
    Ok(())
}
CLI
xberg extract document.pdf
xberg extract document.pdf --format json
xberg extract document.pdf --content-format markdown

Configuration

All languages use the same configuration structure with language-appropriate naming conventions.

Python (snake_case)
from xberg import (
    ExtractInput, extract,
    ExtractionConfig, OcrConfig, TesseractConfig, PdfConfig, ChunkingConfig, OutputFormat,
)

config = ExtractionConfig(
    ocr=OcrConfig(
        backend="tesseract",
        language=["eng"],
        tesseract_config=TesseractConfig(psm=6, enable_table_detection=True),
    ),
    pdf_options=PdfConfig(passwords=["secret123"]),
    chunking=ChunkingConfig(max_characters=1000, overlap=200),
    output_format=OutputFormat("markdown"),
)

result = await extract(ExtractInput(uri="document.pdf"), config)
Node.js (camelCase)
import { extract, type ExtractionConfig } from "@xberg-io/xberg";

const config: ExtractionConfig = {
  ocr: { backend: "tesseract", language: ["eng"] },
  pdfOptions: { passwords: ["secret123"] },
  chunking: { maxCharacters: 1000, overlap: 200 },
  outputFormat: "markdown",
};

const output = await extract({ kind: "uri", uri: "document.pdf" }, config);
Rust (snake_case)
use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig, ChunkingConfig, OutputFormat};

let config = ExtractionConfig {
    ocr: Some(OcrConfig {
        backend: "tesseract".into(),
        language: vec!["eng".to_string()],
        ..Default::default()
    }),
    chunking: Some(ChunkingConfig {
        max_characters: 1000,
        overlap: 200,
        ..Default::default()
    }),
    output_format: OutputFormat::Markdown,
    ..Default::default()
};

let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?;
Config File (TOML)
output_format = "markdown"

[ocr]
backend = "tesseract"
language = "eng"

[chunking]
max_characters = 1000
overlap = 200

[pdf_options]
passwords = ["secret123"]
# CLI: auto-discovers xberg.toml/yaml/yml/json in current/parent directories
xberg extract doc.pdf
# or explicit:
xberg extract doc.pdf --config xberg.toml
xberg extract doc.pdf --config-json '{"ocr":{"backend":"tesseract","language":"deu"}}'

Batch Processing

extract_batch takes a list of ExtractInputs and returns one envelope whose results array holds a document per input (in input order); per-input failures are reported in result.errors.

Python
from xberg import ExtractInput, extract_batch, ExtractionConfig

inputs = [
    ExtractInput(uri="doc1.pdf"),
    ExtractInput(uri="doc2.docx"),
    ExtractInput(uri="doc3.xlsx"),
]
output = await extract_batch(inputs, ExtractionConfig())

for doc in output.results:
    print(f"{len(doc.content)} chars extracted")
Node.js
import { extractBatch } from "@xberg-io/xberg";

const output = await extractBatch([
  { kind: "uri", uri: "doc1.pdf" },
  { kind: "uri", uri: "doc2.docx" },
]);
for (const doc of output.results) {
  console.log(`${doc.content.length} chars`);
}
Rust
use xberg::{extract_batch, ExtractInput, ExtractionConfig};

let config = ExtractionConfig::default();
let inputs = vec![ExtractInput::from_uri("doc1.pdf"), ExtractInput::from_uri("doc2.docx")];
let output = extract_batch(inputs, &config).await?;
CLI
xberg batch *.pdf --format json
xberg batch docs/*.docx --content-format markdown

OCR

OCR runs automatically for images and scanned PDFs. Tesseract is the default backend (native binding, no external install required).

Backends

Select with OcrConfig.backend:

  • tesseract (default): built-in native binding. All Tesseract languages supported.
  • paddleocr ("paddleocr" / "paddle-ocr"): ONNX-based PaddleOCR.
  • vlm: Vision-Language-Model OCR (configure via OcrConfig.vlm_config).

Custom backends can be registered in Python/Node via register_ocr_backend (see Advanced Features).

Language Codes
config = ExtractionConfig(ocr=OcrConfig(language=["eng"]))          # English
config = ExtractionConfig(ocr=OcrConfig(language=["eng", "deu"]))   # Multiple
# The single-string shorthand ("eng+deu") is only accepted in config files / --config-json,
# not in the OcrConfig constructor (Python takes a list, Node takes an array).
Force OCR
config = ExtractionConfig(force_ocr=True)  # OCR even if text is extractable

Result Envelope and Document Fields

extract / extract_batch return an ExtractionResult envelope: results (list of documents), errors (per-input failures), and summary (counts). Per-document fields live on each document in results — bind doc = result.results[0] (Python/Node) or &output.results[0] (Rust) first.

Field Python (doc.) Node.js (doc.) Rust (document.) Description
Text content content content content Extracted text (str/String)
MIME type mime_type mimeType mime_type Input document MIME type
Metadata metadata metadata metadata Document metadata (flat mapping)
Tables tables tables tables Extracted tables with cells + markdown
Languages detected_languages detectedLanguages detected_languages Detected languages (if enabled)
Chunks chunks chunks chunks Text chunks (if chunking enabled)
Images images images images Extracted images (if enabled)
Elements elements elements elements Semantic elements (if element_based format)
Pages pages pages pages Per-page content (if page extraction enabled)
Keywords extracted_keywords extractedKeywords extracted_keywords Extracted keywords (if enabled)

Error Handling

Python

extract / extract_batch raise a plain RuntimeError on failure — the typed XbergError subclasses are not raised by these entry points, so catch RuntimeError. Per-input failures during extract_batch are reported non-fatally in result.errors.

from xberg import ExtractInput, extract, ExtractionConfig

try:
    result = await extract(ExtractInput(uri="file.pdf"), ExtractionConfig())
    for err in result.errors:
        print(f"Per-input error: {err}")
except RuntimeError as e:
    print(f"Extraction failed: {e}")
Node.js

The Node binding throws plain Error objects (it does not export typed error subclasses). Catch with instanceof Error, and inspect output.errors for non-fatal per-input failures.

import { extract } from "@xberg-io/xberg";

try {
  const output = await extract({ kind: "uri", uri: "file.pdf" });
  if (output.errors.length > 0) {
    console.error("Per-input errors:", output.errors);
  }
} catch (e) {
  if (e instanceof Error) {
    console.error(`Extraction failed: ${e.message}`);
  }
}
Rust
use xberg::{extract, ExtractInput, ExtractionConfig, XbergError};

let config = ExtractionConfig::default();
match extract(ExtractInput::from_uri("file.pdf"), &config).await {
    Ok(output) => println!("{}", output.results[0].content),
    Err(XbergError::Parsing { message, .. }) => eprintln!("Parse error: {message}"),
    Err(XbergError::Ocr { message, .. }) => eprintln!("OCR error: {message}"),
    Err(XbergError::UnsupportedFormat(mime)) => eprintln!("Unsupported: {mime}"),
    Err(e) => eprintln!("Error: {e}"),
}

Common Pitfalls

  1. Result is an envelope: extract / extract_batch return ExtractionResult with results, errors, and summary. Per-document fields (content, tables, chunks, …) are on result.results[i], NOT on the top-level return.
  2. Async-only: Python and Node have no sync variants — always await extract(...). Rust extract is async; use #[tokio::main] or an async context.
  3. Build the input: pass an ExtractInput, not a bare path. Use ExtractInput(uri=...) / ExtractInput::from_uri(...) (Python/Rust) or { kind: "uri", uri: "..." } (Node); for bytes use kind="bytes" with bytes/mime_type.
  4. Python ChunkingConfig fields: construct with max_characters and overlap (defaults 1000 / 200); these are also the readable attributes. When passing config as a dict/JSON, the max_chars / max_overlap aliases are also accepted. Node uses maxCharacters / overlap; Rust struct fields are max_characters / overlap.
  5. Python errors: extract / extract_batch raise a plain RuntimeError on failure, not typed XbergError subclasses — catch RuntimeError. Node throws plain Error (no typed error subclasses).
  6. Rust extract signature: extract(input, &config) — the config is a reference. Use &ExtractionConfig::default() for defaults.
  7. CLI --format vs --content-format: --format controls CLI output (text, json, or toon). --content-format controls content rendering (plain, markdown, djot, html, json, or doctags).
  8. Config file field names: Use snake_case in TOML/YAML/JSON config files — [chunking] fields are max_characters and overlap; other fields use names like output_format, pdf_options.

Supported Formats (Summary)

Category Extensions
PDF .pdf
Word .docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pages, .wpd, .wp, .wp5, .wp6, .hwp, .hwpx
Spreadsheets .xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbers
Presentations .pptx, .pptm, .ppt, .pps, .ppsx, .potx, .potm, .pot, .odp, .key
eBooks .epub, .fb2
Images .png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tif, .jp2, .jpg2, .j2c, .j2k, .jpc, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppm, .heic, .heics, .heif, .heifs, .hif, .avif, .avcs, .svg
Markup .html, .htm, .xhtml, .xht, .xml, .kml
Data .json, .geojson, .jsonl, .ndjson, .yaml, .yml, .toml, .csv, .tsv, .dbf, .sqlite, .sqlite3, .db, .gpkg, .gpkx
Text .txt, .adoc, .asciidoc, .vtt, .md, .markdown, .commonmark, .qmd, .rmd, .mdx, .djot, .dj, .doctags, .rst, .org, .rtf
Email .eml, .msg, .pst
Archives .zip, .tar, .tgz, .gz, .7z
Audio/Video .mp3, .mpga, .m4a, .wav, .webm, .mp4, .mpg4, .mp4v, .m4v, .mpeg, .mpg, .mpe, .m1v, .m2v
Academic .bib, .ris, .nbib, .enw, .tex, .latex, .typ, .typst, .jats, .nxml, .ipynb, .docbook, .dbk, .docbook4, .docbook5, .opml

CSL JSON is supported through an explicit MIME type but does not have a registered file extension.

See references/supported-formats.md for the complete format reference with MIME types.

Additional Resources

Detailed reference files for specific topics:

Task-focused sibling skills go deeper than this overview:

  • extracting-with-ocr — OCR backends, language packs, force-OCR, tuning.
  • extracting-tables — layout-aware table detection and table models.
  • chunking — chunk size/overlap, markdown/yaml/semantic chunkers, the chunk command.
  • extracting-keywords — YAKE/RAKE keywords, language detection, the embed command.
  • batch-extraction — the batch command, --file-configs, parallelism, error recovery.
  • picking-a-format — choosing --format / --content-format per consumer.

Full documentation: https://docs.xberg.io GitHub: https://github.com/xberg-io/xberg

1---
2name: xberg
3description: >-
4 Extract text, tables, metadata, and images from 107 document formats
5 (PDF, Office, images, HTML, email, archives, academic) using Xberg.
6 Use when writing code that calls Xberg APIs in Python, Node.js/TypeScript,
7 Rust, or CLI. Covers installation, extraction (sync/async), configuration
8 (OCR, chunking, output format), batch processing, error handling, and plugins.
9license: Elastic-2.0
10metadata:
11 author: xberg-io
12 version: "0.1.0"
13 repository: https://github.com/xberg-io/xberg
14---
15 
16<!--
17AI-RULEZ :: GENERATED FILE — DO NOT EDIT
18Content-Hash: blake3:8d1ddd2c0fd7b6e9473737826178427138c3fac45cedd5b6c3d0b8a243739deb
19Source-Hash: blake3:4fe0d36f1a27b5607297a6d73b66eb41eb2264ca9a2c91ab638ab27a12ebb2cd
20Schema-Version: v1
21-->
22 
23# Xberg Document Extraction
24 
25Xberg is a document intelligence library with a Rust core and bindings for Python, TypeScript/Node.js, Ruby, PHP, Go, Java, C#, Elixir, WebAssembly, Dart, Kotlin Android, Swift, Zig, and C. It extracts text, tables, metadata, and images from 107 formats across 140 unique file extensions and accepts 53 compatibility MIME aliases, including PDF, Office documents, images, HTML, email, archives, and academic formats.
26 
27Use this skill when writing code that:
28 
29- Extracts text or metadata from documents
30- Performs OCR on scanned documents or images
31- Batch-processes multiple files
32- Configures extraction options (output format, chunking, OCR, language detection)
33- Implements custom plugins (post-processors, validators, OCR backends)
34 
35> If the `xberg` MCP server is registered in this session, prefer its tools over shelling out to the CLI — they expose the same extraction surface with structured arguments and results.
36 
37## Installation
38 
39### Python
40 
41```bash
42pip install xberg
43```
44 
45### Node.js
46 
47```bash
48npm install @xberg-io/xberg
49```
50 
51### Rust
52 
53```bash
54cargo add xberg
55```
56 
57```toml
58# Cargo.toml
59[dependencies]
60xberg = { version = "1.1.0", features = ["full"] }
61tokio = { version = "1", features = ["full"] }
62# feature flags: pdf, ocr, chunking, embeddings, language-detection, keywords, api, mcp
63# (or "formats" / "full" aggregates); tokio-runtime is on by default
64```
65 
66### CLI
67 
68```bash
69brew install xberg-io/tap/xberg
70# or run without a persistent install (the CLI proxy package self-installs the binary):
71npx @xberg-io/xberg-cli --help
72uvx --from xberg-cli xberg --help
73# or download a prebuilt binary from the latest GitHub release:
74# https://github.com/xberg-io/xberg/releases/latest
75# or build from source:
76cargo install xberg-cli
77```
78 
79## Quick Start
80 
81The library entry points are `extract(input, config)` and `extract_batch(inputs, config)`. Both return an `ExtractionResult` **envelope** — the extracted document(s) live in `result.results`, and per-document data (`content`, `tables`, `metadata`, …) is on each `result.results[i]`. Python and Node are async-only.
82 
83### Python
84 
85```python
86import asyncio
87from xberg import ExtractInput, extract, ExtractionConfig
88 
89async def main() -> None:
90 result = await extract(ExtractInput(uri="document.pdf"), ExtractionConfig())
91 doc = result.results[0]
92 print(doc.content) # extracted text
93 print(doc.metadata) # document metadata
94 print(doc.tables) # extracted tables
95 
96asyncio.run(main())
97```
98 
99### Node.js
100 
101```typescript
102import { extract } from "@xberg-io/xberg";
103 
104const output = await extract({ kind: "uri", uri: "document.pdf" });
105const doc = output.results[0];
106console.log(doc.content);
107console.log(doc.metadata);
108console.log(doc.tables);
109```
110 
111### Rust
112 
113```rust
114use xberg::{extract, ExtractInput, ExtractionConfig};
115 
116#[tokio::main]
117async fn main() -> xberg::Result<()> {
118 let output = extract(ExtractInput::from_uri("document.pdf"), &ExtractionConfig::default()).await?;
119 println!("{}", output.results[0].content);
120 Ok(())
121}
122```
123 
124### CLI
125 
126```bash
127xberg extract document.pdf
128xberg extract document.pdf --format json
129xberg extract document.pdf --content-format markdown
130```
131 
132## Configuration
133 
134All languages use the same configuration structure with language-appropriate naming conventions.
135 
136### Python (snake_case)
137 
138```python
139from xberg import (
140 ExtractInput, extract,
141 ExtractionConfig, OcrConfig, TesseractConfig, PdfConfig, ChunkingConfig, OutputFormat,
142)
143 
144config = ExtractionConfig(
145 ocr=OcrConfig(
146 backend="tesseract",
147 language=["eng"],
148 tesseract_config=TesseractConfig(psm=6, enable_table_detection=True),
149 ),
150 pdf_options=PdfConfig(passwords=["secret123"]),
151 chunking=ChunkingConfig(max_characters=1000, overlap=200),
152 output_format=OutputFormat("markdown"),
153)
154 
155result = await extract(ExtractInput(uri="document.pdf"), config)
156```
157 
158### Node.js (camelCase)
159 
160```typescript
161import { extract, type ExtractionConfig } from "@xberg-io/xberg";
162 
163const config: ExtractionConfig = {
164 ocr: { backend: "tesseract", language: ["eng"] },
165 pdfOptions: { passwords: ["secret123"] },
166 chunking: { maxCharacters: 1000, overlap: 200 },
167 outputFormat: "markdown",
168};
169 
170const output = await extract({ kind: "uri", uri: "document.pdf" }, config);
171```
172 
173### Rust (snake_case)
174 
175```rust
176use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig, ChunkingConfig, OutputFormat};
177 
178let config = ExtractionConfig {
179 ocr: Some(OcrConfig {
180 backend: "tesseract".into(),
181 language: vec!["eng".to_string()],
182 ..Default::default()
183 }),
184 chunking: Some(ChunkingConfig {
185 max_characters: 1000,
186 overlap: 200,
187 ..Default::default()
188 }),
189 output_format: OutputFormat::Markdown,
190 ..Default::default()
191};
192 
193let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?;
194```
195 
196### Config File (TOML)
197 
198```toml
199output_format = "markdown"
200 
201[ocr]
202backend = "tesseract"
203language = "eng"
204 
205[chunking]
206max_characters = 1000
207overlap = 200
208 
209[pdf_options]
210passwords = ["secret123"]
211```
212 
213```bash
214# CLI: auto-discovers xberg.toml/yaml/yml/json in current/parent directories
215xberg extract doc.pdf
216# or explicit:
217xberg extract doc.pdf --config xberg.toml
218xberg extract doc.pdf --config-json '{"ocr":{"backend":"tesseract","language":"deu"}}'
219```
220 
221## Batch Processing
222 
223`extract_batch` takes a list of `ExtractInput`s and returns one envelope whose `results` array holds a document per input (in input order); per-input failures are reported in `result.errors`.
224 
225### Python
226 
227```python
228from xberg import ExtractInput, extract_batch, ExtractionConfig
229 
230inputs = [
231 ExtractInput(uri="doc1.pdf"),
232 ExtractInput(uri="doc2.docx"),
233 ExtractInput(uri="doc3.xlsx"),
234]
235output = await extract_batch(inputs, ExtractionConfig())
236 
237for doc in output.results:
238 print(f"{len(doc.content)} chars extracted")
239```
240 
241### Node.js
242 
243```typescript
244import { extractBatch } from "@xberg-io/xberg";
245 
246const output = await extractBatch([
247 { kind: "uri", uri: "doc1.pdf" },
248 { kind: "uri", uri: "doc2.docx" },
249]);
250for (const doc of output.results) {
251 console.log(`${doc.content.length} chars`);
252}
253```
254 
255### Rust
256 
257```rust
258use xberg::{extract_batch, ExtractInput, ExtractionConfig};
259 
260let config = ExtractionConfig::default();
261let inputs = vec![ExtractInput::from_uri("doc1.pdf"), ExtractInput::from_uri("doc2.docx")];
262let output = extract_batch(inputs, &config).await?;
263```
264 
265### CLI
266 
267```bash
268xberg batch *.pdf --format json
269xberg batch docs/*.docx --content-format markdown
270```
271 
272## OCR
273 
274OCR runs automatically for images and scanned PDFs. Tesseract is the default backend (native binding, no external install required).
275 
276### Backends
277 
278Select with `OcrConfig.backend`:
279 
280- **tesseract** (default): built-in native binding. All Tesseract languages supported.
281- **paddleocr** (`"paddleocr"` / `"paddle-ocr"`): ONNX-based PaddleOCR.
282- **vlm**: Vision-Language-Model OCR (configure via `OcrConfig.vlm_config`).
283 
284Custom backends can be registered in Python/Node via `register_ocr_backend` (see [Advanced Features](references/advanced-features.md)).
285 
286### Language Codes
287 
288```python
289config = ExtractionConfig(ocr=OcrConfig(language=["eng"])) # English
290config = ExtractionConfig(ocr=OcrConfig(language=["eng", "deu"])) # Multiple
291# The single-string shorthand ("eng+deu") is only accepted in config files / --config-json,
292# not in the OcrConfig constructor (Python takes a list, Node takes an array).
293```
294 
295### Force OCR
296 
297```python
298config = ExtractionConfig(force_ocr=True) # OCR even if text is extractable
299```
300 
301## Result Envelope and Document Fields
302 
303`extract` / `extract_batch` return an `ExtractionResult` envelope: `results` (list of documents), `errors` (per-input failures), and `summary` (counts). Per-document fields live on each document in `results` — bind `doc = result.results[0]` (Python/Node) or `&output.results[0]` (Rust) first.
304 
305| Field | Python (`doc.`) | Node.js (`doc.`) | Rust (`document.`) | Description |
306| ------------ | ---------------------- | --------------------- | ----------------------- | --------------------------------------------- |
307| Text content | `content` | `content` | `content` | Extracted text (str/String) |
308| MIME type | `mime_type` | `mimeType` | `mime_type` | Input document MIME type |
309| Metadata | `metadata` | `metadata` | `metadata` | Document metadata (flat mapping) |
310| Tables | `tables` | `tables` | `tables` | Extracted tables with cells + markdown |
311| Languages | `detected_languages` | `detectedLanguages` | `detected_languages` | Detected languages (if enabled) |
312| Chunks | `chunks` | `chunks` | `chunks` | Text chunks (if chunking enabled) |
313| Images | `images` | `images` | `images` | Extracted images (if enabled) |
314| Elements | `elements` | `elements` | `elements` | Semantic elements (if element_based format) |
315| Pages | `pages` | `pages` | `pages` | Per-page content (if page extraction enabled) |
316| Keywords | `extracted_keywords` | `extractedKeywords` | `extracted_keywords` | Extracted keywords (if enabled) |
317 
318## Error Handling
319 
320### Python
321 
322`extract` / `extract_batch` raise a plain `RuntimeError` on failure — the typed `XbergError` subclasses are not raised by these entry points, so catch `RuntimeError`. Per-input failures during `extract_batch` are reported non-fatally in `result.errors`.
323 
324```python
325from xberg import ExtractInput, extract, ExtractionConfig
326 
327try:
328 result = await extract(ExtractInput(uri="file.pdf"), ExtractionConfig())
329 for err in result.errors:
330 print(f"Per-input error: {err}")
331except RuntimeError as e:
332 print(f"Extraction failed: {e}")
333```
334 
335### Node.js
336 
337The Node binding throws plain `Error` objects (it does not export typed error subclasses). Catch with `instanceof Error`, and inspect `output.errors` for non-fatal per-input failures.
338 
339```typescript
340import { extract } from "@xberg-io/xberg";
341 
342try {
343 const output = await extract({ kind: "uri", uri: "file.pdf" });
344 if (output.errors.length > 0) {
345 console.error("Per-input errors:", output.errors);
346 }
347} catch (e) {
348 if (e instanceof Error) {
349 console.error(`Extraction failed: ${e.message}`);
350 }
351}
352```
353 
354### Rust
355 
356```rust
357use xberg::{extract, ExtractInput, ExtractionConfig, XbergError};
358 
359let config = ExtractionConfig::default();
360match extract(ExtractInput::from_uri("file.pdf"), &config).await {
361 Ok(output) => println!("{}", output.results[0].content),
362 Err(XbergError::Parsing { message, .. }) => eprintln!("Parse error: {message}"),
363 Err(XbergError::Ocr { message, .. }) => eprintln!("OCR error: {message}"),
364 Err(XbergError::UnsupportedFormat(mime)) => eprintln!("Unsupported: {mime}"),
365 Err(e) => eprintln!("Error: {e}"),
366}
367```
368 
369## Common Pitfalls
370 
3711. **Result is an envelope**: `extract` / `extract_batch` return `ExtractionResult` with `results`, `errors`, and `summary`. Per-document fields (`content`, `tables`, `chunks`, …) are on `result.results[i]`, NOT on the top-level return.
3722. **Async-only**: Python and Node have no sync variants — always `await extract(...)`. Rust `extract` is async; use `#[tokio::main]` or an async context.
3733. **Build the input**: pass an `ExtractInput`, not a bare path. Use `ExtractInput(uri=...)` / `ExtractInput::from_uri(...)` (Python/Rust) or `{ kind: "uri", uri: "..." }` (Node); for bytes use `kind="bytes"` with `bytes`/`mime_type`.
3744. **Python ChunkingConfig fields**: construct with `max_characters` and `overlap` (defaults 1000 / 200); these are also the readable attributes. When passing config as a dict/JSON, the `max_chars` / `max_overlap` aliases are also accepted. Node uses `maxCharacters` / `overlap`; Rust struct fields are `max_characters` / `overlap`.
3755. **Python errors**: `extract` / `extract_batch` raise a plain `RuntimeError` on failure, not typed `XbergError` subclasses — catch `RuntimeError`. Node throws plain `Error` (no typed error subclasses).
3766. **Rust extract signature**: `extract(input, &config)` — the config is a reference. Use `&ExtractionConfig::default()` for defaults.
3777. **CLI --format vs --content-format**: `--format` controls CLI output (`text`, `json`, or `toon`). `--content-format` controls content rendering (`plain`, `markdown`, `djot`, `html`, `json`, or `doctags`).
3788. **Config file field names**: Use snake_case in TOML/YAML/JSON config files — `[chunking]` fields are `max_characters` and `overlap`; other fields use names like `output_format`, `pdf_options`.
379 
380## Supported Formats (Summary)
381 
382| Category | Extensions |
383| ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
384| **PDF** | `.pdf` |
385| **Word** | `.docx`, `.docm`, `.doc`, `.dotx`, `.dotm`, `.dot`, `.odt`, `.pages`, `.wpd`, `.wp`, `.wp5`, `.wp6`, `.hwp`, `.hwpx` |
386| **Spreadsheets** | `.xlsx`, `.xlsm`, `.xlsb`, `.xls`, `.xla`, `.xlam`, `.xltm`, `.xltx`, `.xlt`, `.ods`, `.numbers` |
387| **Presentations** | `.pptx`, `.pptm`, `.ppt`, `.pps`, `.ppsx`, `.potx`, `.potm`, `.pot`, `.odp`, `.key` |
388| **eBooks** | `.epub`, `.fb2` |
389| **Images** | `.png`, `.jpg`, `.jpeg`, `.gif`, `.webp`, `.bmp`, `.tiff`, `.tif`, `.jp2`, `.jpg2`, `.j2c`, `.j2k`, `.jpc`, `.jbig2`, `.jb2`, `.pnm`, `.pbm`, `.pgm`, `.ppm`, `.heic`, `.heics`, `.heif`, `.heifs`, `.hif`, `.avif`, `.avcs`, `.svg` |
390| **Markup** | `.html`, `.htm`, `.xhtml`, `.xht`, `.xml`, `.kml` |
391| **Data** | `.json`, `.geojson`, `.jsonl`, `.ndjson`, `.yaml`, `.yml`, `.toml`, `.csv`, `.tsv`, `.dbf`, `.sqlite`, `.sqlite3`, `.db`, `.gpkg`, `.gpkx` |
392| **Text** | `.txt`, `.adoc`, `.asciidoc`, `.vtt`, `.md`, `.markdown`, `.commonmark`, `.qmd`, `.rmd`, `.mdx`, `.djot`, `.dj`, `.doctags`, `.rst`, `.org`, `.rtf` |
393| **Email** | `.eml`, `.msg`, `.pst` |
394| **Archives** | `.zip`, `.tar`, `.tgz`, `.gz`, `.7z` |
395| **Audio/Video** | `.mp3`, `.mpga`, `.m4a`, `.wav`, `.webm`, `.mp4`, `.mpg4`, `.mp4v`, `.m4v`, `.mpeg`, `.mpg`, `.mpe`, `.m1v`, `.m2v` |
396| **Academic** | `.bib`, `.ris`, `.nbib`, `.enw`, `.tex`, `.latex`, `.typ`, `.typst`, `.jats`, `.nxml`, `.ipynb`, `.docbook`, `.dbk`, `.docbook4`, `.docbook5`, `.opml` |
397 
398CSL JSON is supported through an explicit MIME type but does not have a registered file extension.
399 
400See [references/supported-formats.md](references/supported-formats.md) for the complete format reference with MIME types.
401 
402## Additional Resources
403 
404Detailed reference files for specific topics:
405 
406- **[Python API Reference](references/python-api.md)** — All functions, config classes, plugin protocols, exact signatures
407- **[Node.js API Reference](references/nodejs-api.md)** — All functions, TypeScript interfaces, worker pool APIs
408- **[Rust API Reference](references/rust-api.md)** — All functions with feature gates, structs, Cargo.toml examples
409- **[CLI Reference](references/cli-reference.md)** — All commands, flags, config precedence, exit codes
410- **[Configuration Reference](references/configuration.md)** — TOML/YAML/JSON formats, auto-discovery, env vars, full schema
411- **[Supported Formats](references/supported-formats.md)** — Format families, extensions, capabilities, and authoritative discovery commands
412- **[Advanced Features](references/advanced-features.md)** — Plugins, embeddings, MCP server, API server, security limits
413- **[Other Language Bindings](references/other-bindings.md)** — Go, Ruby, Java, C#, PHP, Elixir, WASM, Dart, Kotlin Android, Swift, Zig, C, and Docker
414 
415## Related skills
416 
417Task-focused sibling skills go deeper than this overview:
418 
419- **extracting-with-ocr** — OCR backends, language packs, force-OCR, tuning.
420- **extracting-tables** — layout-aware table detection and table models.
421- **chunking** — chunk size/overlap, markdown/yaml/semantic chunkers, the `chunk` command.
422- **extracting-keywords** — YAKE/RAKE keywords, language detection, the `embed` command.
423- **batch-extraction** — the `batch` command, `--file-configs`, parallelism, error recovery.
424- **picking-a-format** — choosing `--format` / `--content-format` per consumer.
425 
426Full documentation: <https://docs.xberg.io>
427GitHub: <https://github.com/xberg-io/xberg>
428 

Discussion

Alternatives