LLM Wiki — Knowledge Distillation Pattern skill
The foundational knowledge distillation pattern for building and maintaining an AI-powered Obsidian wiki.
by Ar9av·MIT license·★ 3,520 Stars on the repo·GitHub ↗
npx degit Ar9av/obsidian-wiki/.skills/llm-wiki#main ~/.claude/skills/llm-wiki-2Checked ·commit main
Files of LLM Wiki — Knowledge Distillation Pattern
Show the full text699 lines
LLM Wiki — Knowledge Distillation Pattern
You are maintaining a persistent, compounding knowledge base. The wiki is not a chatbot — it is a compiled artifact where knowledge is distilled once and kept current, not re-derived on every query.
Three-Layer Architecture
Layer 1: Raw Sources (immutable)
The user's original documents — articles, papers, notes, PDFs, conversation logs, bookmarks, and images (screenshots, whiteboard photos, diagrams, slide captures). These are never modified by the system. They live wherever the user keeps them (configured via OBSIDIAN_SOURCES_DIR in .env). Images are first-class sources: the ingest skills read them via the Read tool's vision support and treat their interpreted content as inferred unless it's verbatim transcribed text. Image ingestion requires a vision-capable model — models without vision support should skip image sources and report which files were skipped.
Think of raw sources as the "source code" — authoritative but hard to query directly.
Don't confuse this with the in-vault _raw/ staging folder, which is a different thing: a scratch inbox for quick captures and drafts awaiting promotion (see wiki-capture and wiki-ingest). Files there aren't Layer 1 sources, but wiki-ingest still moves rather than deletes them on promotion, since some have no other copy.
Layer 2: The Wiki (LLM-maintained)
A collection of interconnected Obsidian-compatible markdown files organized by category. This is the compiled knowledge — synthesized, cross-referenced, and navigable. Each page has:
- YAML frontmatter (title, category, tags, sources, timestamps)
- Obsidian
[[wikilinks]]connecting related concepts - Clear provenance — every claim traces back to a source
The wiki lives at the path configured via OBSIDIAN_VAULT_PATH in .env.
Layer 3: The Schema (this skill + config)
The rules governing how the wiki is structured — categories, conventions, page templates, and operational workflows. The schema tells the LLM how to maintain the wiki.
Wiki Organization
The vault has two levels of structure: categories (what kind of knowledge) and projects (where the knowledge came from).
Categories
Organize pages into these default categories (customizable in .env):
| Category | Purpose | Example |
|---|---|---|
concepts/ |
Ideas, theories, mental models | concepts/transformer-architecture.md |
entities/ |
People, orgs, tools, projects | entities/andrej-karpathy.md |
skills/ |
How-to knowledge, procedures | skills/fine-tuning-llms.md |
references/ |
Summaries of specific sources; academic papers use the Paper Deep-Dive Template (below) | references/attention-is-all-you-need.md |
synthesis/ |
Cross-cutting analysis across sources | synthesis/scaling-laws-debate.md |
journal/ |
Timestamped observations, session logs | journal/2024-03-15.md |
Projects
Knowledge often belongs to a specific project. The projects/ directory mirrors this:
$OBSIDIAN_VAULT_PATH/
├── projects/
│ ├── my-project/
│ │ ├── my-project.md ← project overview (named after project)
│ │ ├── concepts/ ← project-scoped category pages
│ │ ├── skills/
│ │ └── ...
│ ├── another-project/
│ │ └── ...
│ └── side-project/
│ └── ...
├── concepts/ ← global (cross-project) knowledge
├── entities/
├── skills/
└── ...
When knowledge is project-specific (a debugging technique that only applies to one codebase, a project-specific architecture decision), put it under projects/<project-name>/<category>/.
When knowledge is general (a concept like "React Server Components", a person like "Andrej Karpathy", a widely applicable skill), put it in the global category directory.
Cross-referencing: Project pages should [[wikilink]] to global pages and vice versa. A project's overview page should link to the key concept, skill, and entity pages relevant to that project — whether they live under the project or globally.
Naming rule: The project overview file must be named <project-name>.md, not _project.md. Obsidian's graph view uses the filename as the node label — _project.md makes every project appear as _project in the graph, making it unreadable. So projects/my-project/my-project.md, projects/another-project/another-project.md, etc.
Each project directory has an overview page structured like this:
---
title: >-
My Project
category: project
tags: [ai, web, backend]
source_path: ~/.claude/projects/-Users-name-Documents-projects-my-project
created: 2026-03-01T00:00:00Z
updated: 2026-04-06T00:00:00Z
---
# My Project
One-paragraph summary of what this project is.
## Key Concepts
- [[concepts/some-api]] — used for core functionality
- [[projects/my-project/concepts/main-architecture]] — project-specific architecture
## Related
- [[entities/some-service]] — deployment platform
Special Files
Every wiki has these files at its root:
Write them with
obsidian-wiki memory, never by hand.index.md,log.md,hot.md, and the_meta/tables share one advisory lock and are written atomically; hand edits in a parallel run drop whichever write lands second.obsidian-wiki memory sync <VERB> key=valuedoes all three in one call. The full procedure — verbs, theKey Takeawaysslot that stays yours, the owner profile and todo index — is inreferences/MEMORY.md.
index.md
A content-oriented catalog organized by category. Each entry has a one-line summary and tags. Rebuild this after every ingest operation. Format:
# Wiki Index
## Concepts
- [[transformer-architecture]] — The dominant architecture for sequence modeling ( #ml #architecture)
- [[attention-mechanism]] — Core building block of transformers ( #ml #fundamentals)
## Entities
- [[andrej-karpathy]] — AI researcher, educator, former Tesla AI director ( #person #ml)
Format rule: Add a space after the opening ( and tags.
❌ Don't: description (#tag) — breaks tag parsing
✅ Do: description ( #tag) — proper spacing and tag parsing
log.md
Chronological append-only record tracking every operation. Each entry is parseable:
## Log
- [2024-03-15T10:30:00Z] INGEST source="papers/attention.pdf" pages_updated=12 pages_created=3
- [2024-03-15T11:00:00Z] QUERY query="How do transformers handle long sequences?" result_pages=4
- [2024-03-16T09:00:00Z] LINT issues_found=2 orphans=1 contradictions=1
- [2024-03-17T10:00:00Z] ARCHIVE reason="rebuild" pages=87 destination="_archives/..."
- [2024-03-17T10:05:00Z] REBUILD archived_to="_archives/..." previous_pages=87
.manifest.json
Tracks every source file that has been ingested — path, timestamps, what wiki pages it produced. This is the backbone of the delta system. See the wiki-status skill for the full schema.
The manifest enables:
- Delta computation — what's new or modified since last ingest
- Append mode — only process the delta, not everything
- Audit — which source produced which wiki page
- Staleness detection — source changed but wiki page hasn't been updated
Source key contract (v2). Source keys — the sources keys in .manifest.json, the sources: frontmatter values on pages, and a project's source_repo — MUST be machine-portable. A vault is synced across machines, so a bare absolute path (/Users/..., /home/...) is never a valid stored key. This is the single canonical definition; other skills reference it rather than restating it.
| Where the source lives | Canonical key form | Example |
|---|---|---|
| Inside the vault | vault-relative path — POSIX separators, no leading ./, no .. |
Raw/database/postgres.pdf, Clippings/article.md |
Under $HOME |
home-relative path — starts with ~ |
~/.claude/projects/-Users-name-my-app/abc.jsonl |
| Not a file at all | pseudo-key — any scheme:/:// identifier, treated as opaque |
url:https://example.com/article, agent:claude/<session-id> |
Rules:
- Never store a bare absolute path. Convert before writing, not after.
- Normalize before comparing. Expand
~and environment variables, resolve vault-relative keys against the vault root, and treatscheme:/://pseudo-keys as opaque identifiers. Never compare raw strings without normalizing first. - Identity survives path changes. The same logical source keeps the same key across machines.
- Pseudo-keys are an open namespace. What makes a key a pseudo-key is its shape (
scheme:or://, so it can never be mistaken for a file path), not a fixed list of names. Recommended names:repo:<host/owner/name>for a git project,url:<canonical-url>for a web page,agent:<agent>/<id>for an agent session. A source that is neither in the vault nor under$HOMEstill needs one — do not let it fall back to an absolute path. - Project identity is a repository, not a checkout. In the
projectsblock, identify a project bysource_repo(host/owner/name) rather than a machine path. A machine-specific checkout location, if useful at all, belongs in an optionalsource_cwd_hint(~-relative), never in the identity.
Reading is backward compatible: an existing manifest full of absolute keys keeps working, and scripts/manifest.py migrate <vault> --dry-run converts it to contract v2 (merging collisions, keeping the newest ingested_at). If the vault has moved between machines, its absolute keys are rooted at the old vault path, which matches neither the new vault nor $HOME — pass that old root explicitly with migrate <vault> --from-root <old-vault-root> (repeat the flag if the vault lived at more than one location). The command then reports nothing portable to write — N key(s) kept non-portable rather than claiming success. New writes go through the same normalization, so a skill may pass an absolute path to obsidian-wiki cache-update and still have a portable key land in the manifest.
Recording provenance. When you write a manifest entry, populate pages_created and pages_updated with the vault-relative page paths that source contributed to. This is what makes re-ingestion (when a source changes) able to find the pages to revisit, instead of guessing.
Page Template
When creating a new wiki page, use this structure:
---
title: >-
Page Title
category: concepts
tags: [ml, architecture]
aliases: [alternate name]
relationships:
- target: "[[concepts/related-concept]]"
type: extends
sources: [papers/attention.pdf]
summary: >-
One or two sentences, ≤200 chars, so a reader (or another skill) can preview this page without opening it.
provenance:
extracted: 0.72
inferred: 0.25
ambiguous: 0.03
base_confidence: 0.65
lifecycle: draft
lifecycle_changed: 2024-03-15
tier: supporting
created: 2024-03-15T10:30:00Z
updated: 2024-03-15T10:30:00Z
# Optional. Written only by `obsidian-wiki snapshots set` / `apply` — do not hand-author.
# Quoted wikilinks with |title: Obsidian Properties does not treat Markdown [text](path) as links.
snapshots:
- "[[_raw/_archived/example-clip|example-clip]]"
---
# Page Title
One-paragraph summary of what this page covers.
## Key Ideas
- The source's central claim, paraphrased directly.
- A generalization the source implies but doesn't state outright. ^[inferred]
- A figure two sources disagree on. ^[ambiguous]
Use [[wikilinks]] to connect to related pages.
## Open Questions
Things that are unresolved or need more sources.
## Sources
- [[_raw/_archived/example-clip.md]] — snapshot this page was distilled from
Sources section (required, last body section). Every wiki page ends with ## Sources. Entries must be clickable in Obsidian:
- Local snapshot (raw ingest, dropped PDFs/images, Web Clipper files, anything that landed in
_raw/and was archived):[[_raw/_archived/<filename>]]in the body Sources section (body wikilinks may include.md). YAMLsources:stays origin keys (url:,agent:, repo paths, …), not the archive path. YAMLsnapshots:is separate: after moving a file to_raw/_archived/, runobsidian-wiki snapshots set <page> --archive _raw/_archived/<filename>thenobsidian-wiki cache-updateon that archived path. The CLI writes a List of quoted wikilinks with display text, e.g."[[_raw/_archived/clip|clip]]"(no.mdin the target;|clipis what Properties shows). Do not put Markdown[title](path)insnapshots:— Properties leaves those as unclickable text. The snapshots CLI does not touch the body. Do not link the webpage recorded in clipping frontmatter — that URL is mutable origin metadata. - Fetched URL (
/ingest-urlwith no saved snapshot): a markdown link to the canonical URL, and YAMLsources:asurl:<canonical-url>. - Do not mix those up. A clip of a page is not an ingest-from-URL.
Related wiki pages stay in Related / relationships:, not in Sources.
Parser-safe scalars. Write free-text frontmatter values — at minimum title and summary — with folded scalar syntax (>-) as shown above: a bare scalar containing : (colon + space), #, or quotes breaks YAML parsing, and Obsidian then reports "Invalid properties" and hides the frontmatter. Keep the value indented on the line(s) following title: >- / summary: >-.
Paper Deep-Dive Template
The generic template suits most sources. Academic papers are the exception. For ML/AI/LLM/VLM (and similar) papers landing in references/, the substance lives in the architecture, the equations, and the results table — exactly what a terse "Key Ideas" list flattens away. For these, use the richer template below. This is the one place where "compile, don't retrieve" yields to a thorough, self-contained walkthrough a reader could study instead of the paper.
Obsidian renders the needed primitives natively, so no extra tooling is required: Mermaid fenced diagrams, $$…$$ LaTeX (MathJax), markdown tables, and ![[image]] / ![[paper.pdf#page=N]] embeds.
Use this template only when the source is an academic paper (arXiv/conference) with load-bearing figures or equations. Everything else uses the generic Page Template above. Frontmatter, provenance markers, confidence, lifecycle, and relationships: are unchanged — only the body sections differ.
---
# ...required frontmatter, same as the generic template; category: references...
---
# Paper Title
> [!tldr] One sentence: what's new, plus the headline result.
## Problem & Motivation
What's broken or missing that this paper addresses.
## Method / Architecture
Prose walkthrough. Embed the paper's real architecture figure as the primary
visual (see *Academic papers* in `wiki-ingest` for the PyMuPDF extraction recipe).
Fall back to a Mermaid flowchart only when no figure can be extracted.
![[attachments/<slug>-fig1.png]]
*Figure N (Author Year): one-line caption.*
## Key Equations
The 1–3 core equations as display math, not backtick code:
$$ \mathcal{L} = \mathbb{E}_{x}\!\left[-\log p_\theta(y \mid z)\right] $$
## Results
Headline numbers as a table, not a comma-separated blob — and embed a key
results/motivating figure (scaling plot, benchmark chart, capability collage)
when the paper has one:
| Method | Benchmark | Metric | Cost |
|---|---|---|---|
| Baseline | … | … | … |
| **This paper** | … | … | … |
![[attachments/<slug>-resultsN.png]]
*Figure N (Author Year): one-line caption.*
## Limitations
What the paper concedes or sidesteps. Mark reading-between-the-lines as ^[inferred].
## Related
Typed `[[wikilinks]]` to neighbouring work.
## Sources
- [[_raw/_archived/paper.pdf]] — snapshot distilled (if a local PDF/clip was ingested)
- <https://arxiv.org/abs/XXXX.XXXXX> — only if this ingest fetched the URL and there is no local snapshot
A Mermaid diagram reconstructed from the paper's prose is a synthesis, not a transcription — treat it as ^[inferred] when the interpretation is non-trivial.
Provenance Markers
Every claim on a wiki page has one of three provenance states. Mark them inline so the reader (and future ingest passes) can tell signal from synthesis.
These are framework defaults. A vault's AGENTS.md may add markers or workflow flags. Preserve owner extensions and treat orthogonal workflow flags separately from the extracted/inferred/ambiguous truth-state axis.
| State | Marker | Meaning |
|---|---|---|
| Extracted | (no marker — default) | A paraphrase of something a source actually says. |
| Inferred | ^[inferred] suffix |
An LLM-synthesized claim — a connection, generalization, or implication the source doesn't state directly. |
| Ambiguous | ^[ambiguous] suffix |
Sources disagree, or the source is unclear. |
Example:
- Transformers parallelize across positions, unlike RNNs.
- This is why they scale better on modern hardware. ^[inferred]
- GPT-4 was trained on roughly 13T tokens. ^[ambiguous]
Why this syntax:
^[...]is footnote-adjacent in Obsidian — renders cleanly and never collides with[[wikilinks]].- Inline (suffix) so a single bullet stays a single bullet.
- Default = extracted means existing pages without markers stay valid.
Frontmatter summary: Optionally surface the rough mix at the page level so the user can scan for speculation-heavy pages without reading them:
provenance:
extracted: 0.72 # rough fraction of sentences/bullets with no marker
inferred: 0.25
ambiguous: 0.03
These are best-effort numbers written by the ingest skill at create/update time. wiki-lint recomputes them and flags drift. The block is optional — pages without it are treated as fully extracted by convention.
Typed Relationships
Plain [[wikilinks]] in page bodies carry no semantic weight — they indicate "related to" but not how. The optional relationships: frontmatter block adds typed, directional edges to the knowledge graph.
The relationships: block
relationships:
- target: "[[Transformer Architecture]]"
type: extends
- target: "[[LSTM]]"
type: contradicts
- target: "[[Attention Mechanism]]"
type: implements
Each entry has two required fields:
target— a wikilink (using the same format asOBSIDIAN_LINK_FORMAT) to the related pagetype— one of the allowed semantic types below
Allowed relationship types
The table below is the framework default allowlist. A vault's AGENTS.md may extend it; consumers must use the effective allowlist and preserve owner semantics without coercion.
| Type | Meaning | Example |
|---|---|---|
extends |
This page builds on or generalises the target | GPT extends Transformer Architecture |
implements |
This page is a concrete realisation of the target concept | BERT implements Masked Language Modelling |
contradicts |
This page's claims conflict with or refute the target | Evidence A contradicts Evidence B |
derived_from |
This page is based on or adapted from the target | Fine-tuning is derived from Transfer Learning |
uses |
This page depends on or relies on the target | RAG uses Vector Databases |
replaces |
This page supersedes or deprecates the target | GPT-4 replaces GPT-3 |
related_to |
Catch-all: related but no stronger directional type applies | Concept A is related to Concept B |
Rules
- Optional field — omit the block entirely if no typed relationships are known. Untagged wikilinks remain valid and are treated as
related_tobywiki-export. - Don't duplicate — if
[[foo]]already appears as an inline wikilink, therelationships:entry just enriches it with a type; it is not a second link. - Direction matters — the page declaring the entry is the source;
targetis the destination. Only declare relationships from this page's perspective. - Don't fabricate — only add a typed entry when the source material makes the relationship direction and type clear. When in doubt, use
related_toor omit.
Skills that read relationships:: wiki-export (emits typed edges), cross-linker (writes typed entries when inferring links), wiki-query (surfaces type in answers and walks the typed-edge graph for multi-hop "how is X connected to Y" path queries — bounded BFS over the relationships: adjacency, frontmatter-only).
Confidence and Lifecycle
Every page carries two orthogonal trust signals plus an optional supersession link.
The requiredness and lifecycle values below are framework defaults. A vault's AGENTS.md may extend lifecycle values or make trust fields optional. Validators must apply that effective owner schema while still validating any trust value that is present.
The deterministic lint/trust consumer accepts owner schema through OBSIDIAN_ALLOWED_LIFECYCLES, OBSIDIAN_ALLOWED_RELATIONSHIP_TYPES, OBSIDIAN_REQUIRED_TRUST_FIELDS, and OBSIDIAN_SCHEMA_SOURCE. Resolution precedence is CLI > environment/config > these framework defaults (with lifecycle and relationship extensions additive). Explicit blank or whitespace-only values fail closed; omit the variable to select defaults. wiki-lint/SKILL.md owns the operational invocation contract.
Required fields
base_confidence: 0.65 # [0.0, 1.0] — time-independent quality estimate. Stored once, recomputed on content change.
lifecycle: draft # draft | reviewed | verified | disputed | archived
lifecycle_changed: 2024-03-15 # ISO date of last state transition
# lifecycle_reason: "..." # optional free-text — why the state changed; surfaced by wiki-query
# superseded_by: "[[new-page]]" # wikilink; only when lifecycle=archived
lifecycle_reason and superseded_by are optional. Never fabricate them.
Confidence formula
The formula is a manual base score, not a deterministic URL classifier:
base_confidence = lineage_count_score * 0.5 + source_quality_score * 0.5
lineage_count_score = min(independent_evidence_lineages / 3, 1.0)
source_quality_score = avg(reviewed quality score per independent lineage)
After calculating the raw score, assess whether the evidence covers the page's material claims. Partial coverage may justify keeping or lowering the score; unsupported material claims require source/claim repair before any confidence change. Avoid small score churn without meaningful epistemic change.
Source-quality scores (use the highest-matching bucket):
| Bucket | Score | Examples |
|---|---|---|
paper |
1.0 | arXiv, conference proceedings |
official |
0.9 | *.gov, vendor docs |
documentation |
0.85 | well-maintained third-party docs |
book |
0.8 | books, technical references |
repository |
0.75 | content-addressed repository/code evidence |
blog |
0.55 | personal blogs |
session_transcript |
0.5 | conversation history or completed operation |
forum |
0.4 | Stack Overflow, HN, Reddit, issue-grade reports |
unknown |
0.4 | catch-all/current config |
llm_generated |
0.3 | LLM synthesis or unvalidated memory seed |
An independent evidence lineage is an origin that can corroborate a claim independently. Canonical source IDs remain useful for identity, but identity alone does not prove independence. Collapse dependent evidence before counting:
- files, releases, and commits from one repository → one repository lineage;
- retry/review/fix tasks in one workstream → one task lineage;
- parent/child Kanban records → one task lineage;
- byte-identical memories across profiles → one memory lineage;
- a snapshot plus the mutable source it captures → one lineage;
- aliases or metadata references resolving to one origin → one lineage.
The deterministic wiki-lint path validates _meta/trust-ledger.json; it does not recompute confidence from source strings. New or materially changed pages are marked for manual review. Refresh the ledger only after explicit human approval.
Per-skill defaults (ingest skills compute this automatically):
| Skill | base_confidence | lifecycle |
|---|---|---|
wiki-ingest (URL) |
0.17 + 0.5 × classify(url) |
draft |
wiki-ingest (single doc) |
per-source classifier | draft |
wiki-ingest (multi-doc) |
min(N/3,1)×0.5 + avg_q×0.5 |
draft |
wiki-research |
varies, often 0.85+ | draft |
wiki-capture |
0.42 | draft |
*-history-ingest |
0.42 | draft |
wiki-update |
0.59 | draft |
wiki-synthesize |
min(input_pages.base_confidence) |
draft |
Lifecycle state machine
Five states. stale is not a state — it is a computed overlay: is_stale = (today − updated) > 90 days.
| State | Entered by | Notes |
|---|---|---|
draft |
Any ingest skill on first write | Default for all new pages |
reviewed |
Human edit only | |
verified |
Human edit only | Time alone never demotes verified pages |
disputed |
Manual edit only | Overrides every state except archived in display |
archived |
Manual edit, or ingest skill setting superseded_by |
Terminal |
Only ingest skills set draft. All other transitions require a human editor. Update lifecycle_changed whenever the state changes.
Two edge classes are therefore illegal and are reported by obsidian-wiki lint as illegal_lifecycle_transitions: anything falling back to draft (reviewed|verified|disputed → draft), and any exit from archived (it is terminal — restoring a page is a deliberate delete-and-recreate, not a transition). The check compares against the lifecycle recorded in _meta/trust-ledger.json at the page's last review, so it only sees pages that have been reviewed at least once.
Importance Tiering
The tier: field controls which pages get updated on each ingest pass and their priority in retrieval. As wikis grow, re-reading every page on every ingest wastes tokens — tiering lets ingest and query skills focus effort where it matters most.
Three tiers
| Tier | Meaning | Ingest behavior | Query priority |
|---|---|---|---|
core |
Load-bearing pages — many other pages depend on them (high incoming-link count or bridge position). Always worth updating. | Always update if the source is even marginally relevant | Surfaced first in index and full-read passes |
supporting (default) |
Standard wiki pages with moderate connectivity | Update when the source has clear new claims for this page | Standard priority |
peripheral |
Low-connectivity pages — rarely linked, narrowly scoped | Skip unless the source is primarily about this topic | Last resort; skipped when trimming to context budget |
Assignment rules
- New pages: default to
tier: supporting - Promote to
core: when a page accumulates ≥5 incoming wikilinks or is flagged as a bridge bywiki-statusinsights mode - Demote to
peripheral: when a page has ≤1 incoming link and hasn't been updated in 90+ days - Human override always wins — edit
tier:manually to lock a page at any level - Existing pages without
tier:are treated assupporting(backward compatible — no migration needed)
Who manages tier
wiki-ingestreadstier:to decide whether to update a page on the current passwiki-queryusestier:to order candidates in the index pass and trim to context budgetwiki-statusinsights mode computes graph metrics and suggests tier assignments — it never writes them automaticallywiki-lintflags missingtier:on newly created pages (Phase 2 enforcement, same timeline asbase_confidence)
Retrieval Primitives
Reading the vault is the dominant cost of every read-side skill. Use the cheapest primitive that can answer the question and escalate only when the cheaper one is insufficient. Any skill that needs content from the vault should follow this table rather than jumping straight to full-page reads.
| Need | Primitive | Relative cost |
|---|---|---|
| Does a page exist? What's its title/category/tags? | Read index.md; Grep frontmatter blocks (scope with a pattern that targets ^--- blocks at file heads) |
Cheapest |
| 1–2 sentence preview of a page | Read the summary: field in its frontmatter |
Cheap |
| A specific claim or section inside a page | Grep -A <n> -B <n> "<term>" <file> — returns only the matching lines plus context |
Medium |
| Whole-page content | Read <file> |
Expensive — last resort |
| Relationships across pages | Grep "\[\[.*?\]\]" across the vault, or walk wikilinks from a known page |
Case-by-case |
Search command preference: for shell/file searches, use ripgrep (rg, rg --files) when available; if not, fall back to grep/find. Capitalized Grep/Glob names in these skills are tool-generic primitives for agents that expose those tools.
The rule: escalate only when the cheaper primitive can't answer the question. If you can answer from summary: fields alone, don't read page bodies. If a grepped section with -A 10 -B 2 gives you the claim, don't read the whole page. A 500-line page opened to read 15 lines is 485 lines of wasted tokens.
Why this matters: a 20-page vault lets you get away with full-vault scans. A 200-page vault does not. The primitives above are how the skills framework scales to large vaults without a database.
Skills that consume this table: wiki-query, cross-linker, wiki-lint, wiki-status (insights mode). Any new skill that reads the vault should cite this section rather than reinvent the pattern.
QMD Index Freshness
QMD is an optional search index layered on top of the vault. The markdown vault is the source of truth. Any skill that writes wiki markdown should refresh QMD after the vault write completes, but only when QMD_WIKI_COLLECTION is configured and the local QMD transport is available. If QMD refresh fails, keep the vault changes and report the QMD status separately.
Use the cheapest verification path that proves the new content is visible: qmd update, qmd embed only if vectors are stale or missing, then a targeted qmd get or qmd ls check for one written page or the collection root. Read-only skills should not refresh QMD.
Core Principles
Compile, don't retrieve. The wiki is pre-compiled knowledge. When you ingest a source, update every relevant page — don't just create a summary of the source.
Compound over time. Each ingest should make the wiki smarter, not just bigger. Merge new information into existing pages, resolve contradictions, strengthen cross-references.
Provenance matters. Every claim should trace to a source. When updating a page, note which source prompted the update.
Mark inferences. Default sentences are extracted. Mark synthesized claims with
^[inferred]and contested claims with^[ambiguous]. A wiki that hides its guessing rots silently; one that marks it stays trustworthy.Human curates, LLM maintains. The human decides what sources to add and what questions to ask. The LLM handles the bookkeeping — updating cross-references, maintaining consistency, noting contradictions.
Obsidian is the IDE. The user browses and explores the wiki in Obsidian. Everything must be valid Obsidian markdown with working wikilinks.
Link Format
All internal links connecting wiki pages are controlled by OBSIDIAN_LINK_FORMAT from the resolved config (default: wikilink).
| Setting | Syntax | Example |
|---|---|---|
wikilink (default) |
[[path/to/page]] or [[path/to/page|display text]] |
[[concepts/foo|foo]] |
markdown |
[display text](relative/path.md) |
[foo](../concepts/foo.md) |
Generating markdown-format links
When OBSIDIAN_LINK_FORMAT=markdown:
- Compute the path from the current file's directory to the target
.mdfile using..to climb up as needed. - Use the page title or a natural phrase as display text.
- Always include the
.mdextension.
| Current file | Target | Relative link |
|---|---|---|
index.md |
concepts/foo.md |
[foo](concepts/foo.md) |
concepts/foo.md |
entities/bar.md |
[bar](../entities/bar.md) |
projects/my-project/my-project.md |
concepts/foo.md |
[foo](../../concepts/foo.md) |
projects/my-project/concepts/arch.md |
entities/bar.md |
[bar](../../../entities/bar.md) |
The [[path\|display text]] wikilink form maps to [display text](relative/path.md) in Markdown mode.
Scope: this setting affects only newly written or updated links. Existing vault content is never automatically migrated — users who want to convert old links can run the cross-linker or wiki-lint skill.
Every write skill reads OBSIDIAN_LINK_FORMAT from config before generating links and applies the correct format.
Config Resolution Protocol
All skills must resolve config using this algorithm — do not hard-code .env or the global config path directly. This ensures single-vault, multi-vault, project-local, and VPS setups all work correctly.
Global config directory
The global config directory is XDG-style: $XDG_CONFIG_HOME/obsidian-wiki (default ~/.config/obsidian-wiki). Installs that already have a ~/.obsidian-wiki directory keep using it — so an existing setup never breaks — but any new install lands under the XDG path. Resolve it with:
obsidian_wiki_config_dir() {
local xdg_dir="${XDG_CONFIG_HOME:-$HOME/.config}/obsidian-wiki"
local legacy_dir="$HOME/.obsidian-wiki"
if [[ -d "$legacy_dir" && ! -e "$xdg_dir" ]]; then
echo "$legacy_dir"
else
echo "$xdg_dir"
fi
}
Everywhere below, "the global config dir" means $(obsidian_wiki_config_dir), and "the global config" means $(obsidian_wiki_config_dir)/config.
Resolution order
- Inline vault override (
@name) — if the user's request contains an@<name>token (e.g.@work save this,query @personal about X), resolve<global config dir>/config.<name>directly and use itsOBSIDIAN_VAULT_PATH. This overrides both the CWD.envwalk-up and the active symlink, and applies to that invocation only — never runln -sfor otherwise change the active vault for an@namerequest. If<global config dir>/config.<name>doesn't exist, tell the user it doesn't exist and list the available vaults (thewiki-switchList logic), then stop — do not silently fall back to the default. The@nameis a routing directive, not content: strip it out before treating the rest of the request as the actual instruction or page text. - Walk up from CWD — look for a
.envfile in the current directory, then each parent, up to$HOME. Stop at the first.envthat containsOBSIDIAN_VAULT_PATH. If its value is empty, stop there too: tell the user which.envblocked resolution (a blank line copied from.env.exampledoes this) instead of falling through to the global config. - Global config — if no local
.envfound, read the global config ($(obsidian_wiki_config_dir)/config). - Prompt setup — if neither exists, tell the user: "No config found. Run
wiki-setupto initialize your wiki."
@name is a per-invocation override — it targets one vault for one request. /wiki-switch <name> is the persistent default — it re-points the active symlink for all future requests. Use @name to touch the other vault from anywhere without disturbing your default ("brain") vault.
find_config() {
# $1 = parsed @name from the request, if any (else empty)
local config_dir
config_dir="$(obsidian_wiki_config_dir)"
if [[ -n "$1" ]]; then
[[ -f "$config_dir/config.$1" ]] && { echo "$config_dir/config.$1"; return; }
echo ""; return # named vault missing → caller reports + lists, no fallback
fi
dir="$PWD"
while [[ "$dir" != "$HOME" && "$dir" != "/" ]]; do
[[ -f "$dir/.env" ]] && grep -q "OBSIDIAN_VAULT_PATH" "$dir/.env" && { echo "$dir/.env"; return; }
dir="$(dirname "$dir")"
done
[[ -f "$config_dir/config" ]] && { echo "$config_dir/config"; return; }
echo ""
}
Vault-scoped state
Skills that write runtime state (e.g. daily-update) must scope that state to the resolved vault, not to a global path. Use:
VAULT_ID=$(echo "$OBSIDIAN_VAULT_PATH" | md5sum 2>/dev/null || md5 -q - <<< "$OBSIDIAN_VAULT_PATH" | cut -c1-8)
STATE_DIR="$(obsidian_wiki_config_dir)/state/$VAULT_ID"
Standard "Before You Start" block
Every skill's setup section should read:
Resolve config — follow the Config Resolution Protocol in
llm-wiki/SKILL.md. Honor an inline@nameoverride first, then walk up from CWD for.env, fall back to the global config, else prompt setup. This givesOBSIDIAN_VAULT_PATHand any tool-specific path overrides.
Writing Profile Resolution
Before drafting or rewriting natural-language Markdown, resolve the global config directory with the XDG/legacy algorithm above, then read <global config dir>/WRITING.md when it exists. A missing or empty WRITING.md means there are no custom writing preferences. If that optional read fails, warn and continue with the default framework guidance.
The effective precedence is framework invariants > current task/skill requirements > current project AGENTS.md > vault AGENTS.md > global WRITING.md. Framework invariants include schema, provenance, and safety; operation-specific requirements remain authoritative for the current task. Unspecified project and vault rules are inherited from less-specific layers, and more specific same-topic rules win.
Writing preferences apply only to newly drafted or rewritten natural-language fields and body content. This includes natural-language title and summary values in YAML frontmatter, but preferences cannot alter YAML syntax, required keys, structure, types, or machine-generated fields. JSON, structured logs, and pass-through content remain unchanged and retain their required formats and source fidelity.
Environment Variables
The wiki is configured through environment variables (see .env.example). The only required variable is the vault path — everything else has sensible defaults.
OBSIDIAN_VAULT_PATH— Where the wiki lives (required)OBSIDIAN_SOURCES_DIR— Where raw source documents areOBSIDIAN_CATEGORIES— Comma-separated list of categoriesWIKI_SKIP_PROJECTS— Comma-separated substrings; any project dir whose name contains one is excluded from history ingest (scan + delta + manifest). See the "Project Scoping" step in the history-ingest skills.CLAUDE_HISTORY_PATH— Where to find Claude conversation dataCODEX_HISTORY_PATH— Where to find Codex session dataHERMES_HOME— Where to find Hermes agent dataOPENCLAW_HOME— Where to find OpenClaw dataCOPILOT_HISTORY_PATH— Where to find Copilot session dataOBSIDIAN_LINK_FORMAT— Internal link syntax:wikilink(default) ormarkdownWIKI_TOKEN_WARN_THRESHOLD— Emit a warning inwiki-statuswhen the full-wiki token estimate exceeds this value (default:100000). Set to0to disable. Seewiki-statusfor the token footprint report.WIKI_STAGED_WRITES— Whentrue, all LLM-written pages go to_staging/<category>/for human review before promotion. Seewiki-setupandwiki-stage-commitfor details.CODE_UNDERSTANDING_BACKEND— how wiki-update understands a project before distilling:auto(CodeGraph when available, else builtin ast-extract + rg; default),builtin, orcodegraph(explicitly require; warn/error if unavailable).CODE_UNDERSTANDING_CODEGRAPH_BIN— optional path to the codegraph binary when it isn't on PATH.CODE_UNDERSTANDING_CODEGRAPH_BIN— optional path to the codegraph binary when it isn't on PATH. Both resolve likeOBSIDIAN_VAULT_PATH: a real environment variable wins (empty counts as unset), then the nearest.envwalking up from the project directory, then the global config ($(obsidian_wiki_config_dir)/config), then the default.OBSIDIAN_MAX_PAGES_PER_INGEST— cap on pages created/updated perwiki-ingestrun (default:15). Seewiki-ingest, Step 4.LINT_SCHEDULE— how oftendaily-updatealso runswiki-lint:daily|weekly(default) |manual. Seedaily-update, Step 4a.
No API keys are needed — the agent running these skills already has LLM access built in.
Modes of Operation
The wiki supports three ingest modes:
| Mode | When to use | What happens |
|---|---|---|
| Append | Small delta, incremental updates | Compute delta via manifest, ingest only new/modified sources |
| Rebuild | Major drift, fresh start needed | Archive current wiki to _archives/, clear, reprocess all sources |
| Restore | Need to go back | Bring back a previous archive |
Use wiki-status to see the delta and get a recommendation. Use wiki-rebuild for archive/rebuild/restore operations.
Reference
For details on specific operations, see the companion skills:
- wiki-status — Audit what's ingested, compute delta, recommend append vs rebuild
- wiki-rebuild — Archive current wiki, rebuild from scratch, or restore from archive
- wiki-ingest — Distill source documents into wiki pages and raw text/chat/log data
- claude-history-ingest — Ingest Claude conversation history
- codex-history-ingest — Ingest Codex CLI session history
- wiki-query — Answer questions against the wiki
- wiki-lint — Audit and maintain wiki health
- wiki-setup — Initialize a new vault
| 1 | |
| 2 | name llm-wiki |
| 3 | description > |
| 4 | The foundational knowledge distillation pattern for building and maintaining an AI-powered Obsidian wiki. |
| 5 | Based on Andrej Karpathy's LLM Wiki architecture. Use this skill whenever the user wants to understand the |
| 6 | wiki pattern, set up a new knowledge base, or needs guidance on the three-layer architecture (raw sources → |
| 7 | wiki → schema). Also use when discussing knowledge management strategy, wiki structure decisions, or how |
| 8 | to organize distilled knowledge. This is the "theory" skill — other skills handle specific operations |
| 9 | (ingesting, querying, linting). |
| 10 | |
| 11 | |
| 12 | # LLM Wiki — Knowledge Distillation Pattern |
| 13 | |
| 14 | You are maintaining a persistent, compounding knowledge base. The wiki is not a chatbot — it is a **compiled artifact** where knowledge is distilled once and kept current, not re-derived on every query. |
| 15 | |
| 16 | ## Three-Layer Architecture |
| 17 | |
| 18 | ### Layer 1: Raw Sources (immutable) |
| 19 | |
| 20 | The user's original documents — articles, papers, notes, PDFs, conversation logs, bookmarks, **and images** (screenshots, whiteboard photos, diagrams, slide captures). These are never modified by the system. They live wherever the user keeps them (configured via `OBSIDIAN_SOURCES_DIR` in `.env`). Images are first-class sources: the ingest skills read them via the Read tool's vision support and treat their interpreted content as inferred unless it's verbatim transcribed text. Image ingestion requires a vision-capable model — models without vision support should skip image sources and report which files were skipped. |
| 21 | |
| 22 | Think of raw sources as the "source code" — authoritative but hard to query directly. |
| 23 | |
| 24 | Don't confuse this with the in-vault `_raw/` staging folder, which is a different thing: a scratch inbox for quick captures and drafts awaiting promotion (see `wiki-capture` and `wiki-ingest`). Files there aren't Layer 1 sources, but `wiki-ingest` still moves rather than deletes them on promotion, since some have no other copy. |
| 25 | |
| 26 | ### Layer 2: The Wiki (LLM-maintained) |
| 27 | |
| 28 | A collection of interconnected Obsidian-compatible markdown files organized by category. This is the compiled knowledge — synthesized, cross-referenced, and navigable. Each page has: |
| 29 | |
| 30 | YAML frontmatter (title, category, tags, sources, timestamps) |
| 31 | Obsidian `[[wikilinks]]` connecting related concepts |
| 32 | Clear provenance — every claim traces back to a source |
| 33 | |
| 34 | The wiki lives at the path configured via `OBSIDIAN_VAULT_PATH` in `.env`. |
| 35 | |
| 36 | ### Layer 3: The Schema (this skill + config) |
| 37 | |
| 38 | The rules governing how the wiki is structured — categories, conventions, page templates, and operational workflows. The schema tells the LLM *how* to maintain the wiki. |
| 39 | |
| 40 | ## Wiki Organization |
| 41 | |
| 42 | The vault has two levels of structure: **categories** (what kind of knowledge) and **projects** (where the knowledge came from). |
| 43 | |
| 44 | ### Categories |
| 45 | |
| 46 | Organize pages into these default categories (customizable in `.env`): |
| 47 | |
| 48 | | Category | Purpose | Example | |
| 49 | |---|---|---| |
| 50 | | `concepts/` | Ideas, theories, mental models | `concepts/transformer-architecture.md` | |
| 51 | | `entities/` | People, orgs, tools, projects | `entities/andrej-karpathy.md` | |
| 52 | | `skills/` | How-to knowledge, procedures | `skills/fine-tuning-llms.md` | |
| 53 | | `references/` | Summaries of specific sources; academic papers use the Paper Deep-Dive Template (below) | `references/attention-is-all-you-need.md` | |
| 54 | | `synthesis/` | Cross-cutting analysis across sources | `synthesis/scaling-laws-debate.md` | |
| 55 | | `journal/` | Timestamped observations, session logs | `journal/2024-03-15.md` | |
| 56 | |
| 57 | ### Projects |
| 58 | |
| 59 | Knowledge often belongs to a specific project. The `projects/` directory mirrors this: |
| 60 | |
| 61 | |
| 62 | $OBSIDIAN_VAULT_PATH/ |
| 63 | ├── projects/ |
| 64 | │ ├── my-project/ |
| 65 | │ │ ├── my-project.md ← project overview (named after project) |
| 66 | │ │ ├── concepts/ ← project-scoped category pages |
| 67 | │ │ ├── skills/ |
| 68 | │ │ └── ... |
| 69 | │ ├── another-project/ |
| 70 | │ │ └── ... |
| 71 | │ └── side-project/ |
| 72 | │ └── ... |
| 73 | ├── concepts/ ← global (cross-project) knowledge |
| 74 | ├── entities/ |
| 75 | ├── skills/ |
| 76 | └── ... |
| 77 | |
| 78 | |
| 79 | **When knowledge is project-specific** (a debugging technique that only applies to one codebase, a project-specific architecture decision), put it under `projects/<project-name>/<category>/`. |
| 80 | |
| 81 | **When knowledge is general** (a concept like "React Server Components", a person like "Andrej Karpathy", a widely applicable skill), put it in the global category directory. |
| 82 | |
| 83 | **Cross-referencing:** Project pages should `[[wikilink]]` to global pages and vice versa. A project's overview page should link to the key concept, skill, and entity pages relevant to that project — whether they live under the project or globally. |
| 84 | |
| 85 | **Naming rule:** The project overview file must be named `<project-name>.md`, not `_project.md`. Obsidian's graph view uses the filename as the node label — `_project.md` makes every project appear as `_project` in the graph, making it unreadable. So `projects/my-project/my-project.md`, `projects/another-project/another-project.md`, etc. |
| 86 | |
| 87 | Each project directory has an overview page structured like this: |
| 88 | |
| 89 | |
| 90 | |
| 91 | title: >- |
| 92 | My Project |
| 93 | category: project |
| 94 | tags: [ai, web, backend] |
| 95 | source_path: ~/.claude/projects/-Users-name-Documents-projects-my-project |
| 96 | created: 2026-03-01T00:00:00Z |
| 97 | updated: 2026-04-06T00:00:00Z |
| 98 | |
| 99 | |
| 100 | # My Project |
| 101 | |
| 102 | One-paragraph summary of what this project is. |
| 103 | |
| 104 | ## Key Concepts |
| 105 | - [[concepts/some-api]] — used for core functionality |
| 106 | - [[projects/my-project/concepts/main-architecture]] — project-specific architecture |
| 107 | |
| 108 | ## Related |
| 109 | - [[entities/some-service]] — deployment platform |
| 110 | |
| 111 | |
| 112 | ## Special Files |
| 113 | |
| 114 | Every wiki has these files at its root: |
| 115 | |
| 116 | > **Write them with `obsidian-wiki memory`, never by hand.** `index.md`, |
| 117 | > `log.md`, `hot.md`, and the `_meta/` tables share one advisory lock and are |
| 118 | > written atomically; hand edits in a parallel run drop whichever write lands |
| 119 | > second. `obsidian-wiki memory sync <VERB> key=value` does all three |
| 120 | > in one call. The full procedure — verbs, the `Key Takeaways` slot that stays |
| 121 | > yours, the owner profile and todo index — is in |
| 122 | > [`references/MEMORY.md`]. |
| 123 | |
| 124 | ### `index.md` |
| 125 | A content-oriented catalog organized by category. Each entry has a one-line summary and tags. Rebuild this after every ingest operation. Format: |
| 126 | |
| 127 | |
| 128 | # Wiki Index |
| 129 | |
| 130 | ## Concepts |
| 131 | - [[transformer-architecture]] — The dominant architecture for sequence modeling ( #ml #architecture) |
| 132 | - [[attention-mechanism]] — Core building block of transformers ( #ml #fundamentals) |
| 133 | |
| 134 | ## Entities |
| 135 | - [[andrej-karpathy]] — AI researcher, educator, former Tesla AI director ( #person #ml) |
| 136 | |
| 137 | **Format rule**: Add a space after the opening `(` and tags. |
| 138 | ❌ Don't: `description (#tag)` — breaks tag parsing |
| 139 | ✅ Do: `description ( #tag)` — proper spacing and tag parsing |
| 140 | |
| 141 | ### `log.md` |
| 142 | Chronological append-only record tracking every operation. Each entry is parseable: |
| 143 | |
| 144 | |
| 145 | ## Log |
| 146 | |
| 147 | - [2024-03-15T10:30:00Z] INGEST source="papers/attention.pdf" pages_updated=12 pages_created=3 |
| 148 | - [2024-03-15T11:00:00Z] QUERY query="How do transformers handle long sequences?" result_pages=4 |
| 149 | - [2024-03-16T09:00:00Z] LINT issues_found=2 orphans=1 contradictions=1 |
| 150 | - [2024-03-17T10:00:00Z] ARCHIVE reason="rebuild" pages=87 destination="_archives/..." |
| 151 | - [2024-03-17T10:05:00Z] REBUILD archived_to="_archives/..." previous_pages=87 |
| 152 | |
| 153 | |
| 154 | ### `.manifest.json` |
| 155 | Tracks every source file that has been ingested — path, timestamps, what wiki pages it produced. This is the backbone of the delta system. See the `wiki-status` skill for the full schema. |
| 156 | |
| 157 | The manifest enables: |
| 158 | **Delta computation** — what's new or modified since last ingest |
| 159 | **Append mode** — only process the delta, not everything |
| 160 | **Audit** — which source produced which wiki page |
| 161 | **Staleness detection** — source changed but wiki page hasn't been updated |
| 162 | |
| 163 | **Source key contract (v2).** Source keys — the `sources` keys in `.manifest.json`, the `sources:` frontmatter values on pages, and a project's `source_repo` — MUST be machine-portable. A vault is synced across machines, so a bare absolute path (`/Users/...`, `/home/...`) is never a valid stored key. This is the single canonical definition; other skills reference it rather than restating it. |
| 164 | |
| 165 | | Where the source lives | Canonical key form | Example | |
| 166 | |---|---|---| |
| 167 | | Inside the vault | **vault-relative path** — POSIX separators, no leading `./`, no `..` | `Raw/database/postgres.pdf`, `Clippings/article.md` | |
| 168 | | Under `$HOME` | **home-relative path** — starts with `~` | `~/.claude/projects/-Users-name-my-app/abc.jsonl` | |
| 169 | | Not a file at all | **pseudo-key** — any `scheme:`/`://` identifier, treated as opaque | `url:https://example.com/article`, `agent:claude/<session-id>` | |
| 170 | |
| 171 | Rules: |
| 172 | |
| 173 | **Never store a bare absolute path.** Convert before writing, not after. |
| 174 | **Normalize before comparing.** Expand `~` and environment variables, resolve vault-relative keys against the vault root, and treat `scheme:`/`://` pseudo-keys as opaque identifiers. Never compare raw strings without normalizing first. |
| 175 | **Identity survives path changes.** The same logical source keeps the same key across machines. |
| 176 | **Pseudo-keys are an open namespace.** What makes a key a pseudo-key is its shape (`scheme:` or `://`, so it can never be mistaken for a file path), not a fixed list of names. Recommended names: `repo:<host/owner/name>` for a git project, `url:<canonical-url>` for a web page, `agent:<agent>/<id>` for an agent session. A source that is neither in the vault nor under `$HOME` still needs one — do not let it fall back to an absolute path. |
| 177 | **Project identity is a repository, not a checkout.** In the `projects` block, identify a project by `source_repo` (`host/owner/name`) rather than a machine path. A machine-specific checkout location, if useful at all, belongs in an optional `source_cwd_hint` (`~`-relative), never in the identity. |
| 178 | |
| 179 | Reading is backward compatible: an existing manifest full of absolute keys keeps working, and `scripts/manifest.py migrate <vault> --dry-run` converts it to contract v2 (merging collisions, keeping the newest `ingested_at`). **If the vault has moved between machines**, its absolute keys are rooted at the *old* vault path, which matches neither the new vault nor `$HOME` — pass that old root explicitly with `migrate <vault> --from-root <old-vault-root>` (repeat the flag if the vault lived at more than one location). The command then reports `nothing portable to write — N key(s) kept non-portable` rather than claiming success. New writes go through the same normalization, so a skill may pass an absolute path to `obsidian-wiki cache-update` and still have a portable key land in the manifest. |
| 180 | |
| 181 | **Recording provenance.** When you write a manifest entry, populate `pages_created` and `pages_updated` with the vault-relative page paths that source contributed to. This is what makes re-ingestion (when a source changes) able to find the pages to revisit, instead of guessing. |
| 182 | |
| 183 | ## Page Template |
| 184 | |
| 185 | When creating a new wiki page, use this structure: |
| 186 | |
| 187 | |
| 188 | |
| 189 | title: >- |
| 190 | Page Title |
| 191 | category: concepts |
| 192 | tags: [ml, architecture] |
| 193 | aliases: [alternate name] |
| 194 | relationships: |
| 195 | - target: "[[concepts/related-concept]]" |
| 196 | type: extends |
| 197 | sources: [papers/attention.pdf] |
| 198 | summary: >- |
| 199 | One or two sentences, ≤200 chars, so a reader (or another skill) can preview this page without opening it. |
| 200 | provenance: |
| 201 | extracted: 0.72 |
| 202 | inferred: 0.25 |
| 203 | ambiguous: 0.03 |
| 204 | base_confidence: 0.65 |
| 205 | lifecycle: draft |
| 206 | lifecycle_changed: 2024-03-15 |
| 207 | tier: supporting |
| 208 | created: 2024-03-15T10:30:00Z |
| 209 | updated: 2024-03-15T10:30:00Z |
| 210 | # Optional. Written only by `obsidian-wiki snapshots set` / `apply` — do not hand-author. |
| 211 | # Quoted wikilinks with |title: Obsidian Properties does not treat Markdown [text](path) as links. |
| 212 | snapshots: |
| 213 | - "[[_raw/_archived/example-clip|example-clip]]" |
| 214 | |
| 215 | |
| 216 | # Page Title |
| 217 | |
| 218 | One-paragraph summary of what this page covers. |
| 219 | |
| 220 | ## Key Ideas |
| 221 | |
| 222 | - The source's central claim, paraphrased directly. |
| 223 | - A generalization the source implies but doesn't state outright. ^[inferred] |
| 224 | - A figure two sources disagree on. ^[ambiguous] |
| 225 | |
| 226 | Use [[wikilinks]] to connect to related pages. |
| 227 | |
| 228 | ## Open Questions |
| 229 | |
| 230 | Things that are unresolved or need more sources. |
| 231 | |
| 232 | ## Sources |
| 233 | |
| 234 | - [[_raw/_archived/example-clip.md]] — snapshot this page was distilled from |
| 235 | |
| 236 | |
| 237 | **Sources section (required, last body section).** Every wiki page ends with `## Sources`. Entries must be clickable in Obsidian: |
| 238 | |
| 239 | **Local snapshot** (raw ingest, dropped PDFs/images, Web Clipper files, anything that landed in `_raw/` and was archived): `[[_raw/_archived/<filename>]]` in the body **Sources** section (body wikilinks may include `.md`). YAML `sources:` stays origin keys (`url:`, `agent:`, repo paths, …), **not** the archive path. YAML `snapshots:` is separate: after moving a file to `_raw/_archived/`, run `obsidian-wiki snapshots set <page> --archive _raw/_archived/<filename>` then `obsidian-wiki cache-update` on that **archived** path. The CLI writes a **List** of quoted wikilinks with display text, e.g. `"[[_raw/_archived/clip|clip]]"` (no `.md` in the target; `|clip` is what Properties shows). Do **not** put Markdown `[title]` in `snapshots:` — Properties leaves those as unclickable text. The snapshots CLI does not touch the body. Do not link the webpage recorded in clipping frontmatter — that URL is mutable origin metadata. |
| 240 | **Fetched URL** (`/ingest-url` with no saved snapshot): a markdown link to the canonical URL, and YAML `sources:` as `url:<canonical-url>`. |
| 241 | Do not mix those up. A clip of a page is not an ingest-from-URL. |
| 242 | |
| 243 | Related wiki pages stay in **Related** / `relationships:`, not in Sources. |
| 244 | |
| 245 | **Parser-safe scalars.** Write free-text frontmatter values — at minimum `title` and `summary` — with folded scalar syntax (`>-`) as shown above: a bare scalar containing `: ` (colon + space), `#`, or quotes breaks YAML parsing, and Obsidian then reports "Invalid properties" and hides the frontmatter. Keep the value indented on the line(s) following `title: >-` / `summary: >-`. |
| 246 | |
| 247 | ## Paper Deep-Dive Template |
| 248 | |
| 249 | The generic template suits most sources. **Academic papers are the exception.** For ML/AI/LLM/VLM (and similar) papers landing in `references/`, the substance lives in the architecture, the equations, and the results table — exactly what a terse "Key Ideas" list flattens away. For these, use the richer template below. This is the one place where *"compile, don't retrieve"* yields to a thorough, self-contained walkthrough a reader could study instead of the paper. |
| 250 | |
| 251 | Obsidian renders the needed primitives natively, so no extra tooling is required: Mermaid fenced diagrams, `$$…$$` LaTeX (MathJax), markdown tables, and `![[image]]` / `![[paper.pdf#page=N]]` embeds. |
| 252 | |
| 253 | Use this template only when the source is an academic paper (arXiv/conference) with load-bearing figures or equations. Everything else uses the generic Page Template above. Frontmatter, provenance markers, confidence, lifecycle, and `relationships:` are unchanged — only the body sections differ. |
| 254 | |
| 255 | |
| 256 | |
| 257 | # ...required frontmatter, same as the generic template; category: references... |
| 258 | |
| 259 | |
| 260 | # Paper Title |
| 261 | |
| 262 | > [!tldr] One sentence: what's new, plus the headline result. |
| 263 | |
| 264 | ## Problem & Motivation |
| 265 | |
| 266 | What's broken or missing that this paper addresses. |
| 267 | |
| 268 | ## Method / Architecture |
| 269 | |
| 270 | Prose walkthrough. Embed the paper's real architecture figure as the primary |
| 271 | visual (see *Academic papers* in `wiki-ingest` for the PyMuPDF extraction recipe). |
| 272 | Fall back to a Mermaid flowchart only when no figure can be extracted. |
| 273 | |
| 274 | ![[attachments/<slug>-fig1.png]] |
| 275 | *Figure N (Author Year): one-line caption.* |
| 276 | |
| 277 | ## Key Equations |
| 278 | |
| 279 | The 1–3 core equations as display math, not backtick code: |
| 280 | |
| 281 | $$ \mathcal{L} = \mathbb{E}_{x}\!\left[-\log p_\theta(y \mid z)\right] $$ |
| 282 | |
| 283 | ## Results |
| 284 | |
| 285 | Headline numbers as a table, not a comma-separated blob — and embed a key |
| 286 | results/motivating figure (scaling plot, benchmark chart, capability collage) |
| 287 | when the paper has one: |
| 288 | |
| 289 | | Method | Benchmark | Metric | Cost | |
| 290 | |---|---|---|---| |
| 291 | | Baseline | … | … | … | |
| 292 | | **This paper** | … | … | … | |
| 293 | |
| 294 | ![[attachments/<slug>-resultsN.png]] |
| 295 | *Figure N (Author Year): one-line caption.* |
| 296 | |
| 297 | ## Limitations |
| 298 | |
| 299 | What the paper concedes or sidesteps. Mark reading-between-the-lines as ^[inferred]. |
| 300 | |
| 301 | ## Related |
| 302 | |
| 303 | Typed `[[wikilinks]]` to neighbouring work. |
| 304 | |
| 305 | ## Sources |
| 306 | |
| 307 | - [[_raw/_archived/paper.pdf]] — snapshot distilled (if a local PDF/clip was ingested) |
| 308 | - <https://arxiv.org/abs/XXXX.XXXXX> — only if this ingest fetched the URL and there is no local snapshot |
| 309 | |
| 310 | |
| 311 | A Mermaid diagram reconstructed from the paper's prose is a synthesis, not a transcription — treat it as `^[inferred]` when the interpretation is non-trivial. |
| 312 | |
| 313 | ## Provenance Markers |
| 314 | |
| 315 | Every claim on a wiki page has one of three provenance states. Mark them inline so the reader (and future ingest passes) can tell signal from synthesis. |
| 316 | |
| 317 | These are framework defaults. A vault's `AGENTS.md` may add markers or workflow flags. Preserve owner extensions and treat orthogonal workflow flags separately from the extracted/inferred/ambiguous truth-state axis. |
| 318 | |
| 319 | | State | Marker | Meaning | |
| 320 | |---|---|---| |
| 321 | | **Extracted** | *(no marker — default)* | A paraphrase of something a source actually says. | |
| 322 | | **Inferred** | `^[inferred]` suffix | An LLM-synthesized claim — a connection, generalization, or implication the source doesn't state directly. | |
| 323 | | **Ambiguous** | `^[ambiguous]` suffix | Sources disagree, or the source is unclear. | |
| 324 | |
| 325 | Example: |
| 326 | |
| 327 | |
| 328 | - Transformers parallelize across positions, unlike RNNs. |
| 329 | - This is why they scale better on modern hardware. ^[inferred] |
| 330 | - GPT-4 was trained on roughly 13T tokens. ^[ambiguous] |
| 331 | |
| 332 | |
| 333 | **Why this syntax:** |
| 334 | `^[...]` is footnote-adjacent in Obsidian — renders cleanly and never collides with `[[wikilinks]]`. |
| 335 | Inline (suffix) so a single bullet stays a single bullet. |
| 336 | Default = extracted means existing pages without markers stay valid. |
| 337 | |
| 338 | **Frontmatter summary:** Optionally surface the rough mix at the page level so the user can scan for speculation-heavy pages without reading them: |
| 339 | |
| 340 | |
| 341 | provenance: |
| 342 | extracted: 0.72 # rough fraction of sentences/bullets with no marker |
| 343 | inferred: 0.25 |
| 344 | ambiguous: 0.03 |
| 345 | |
| 346 | |
| 347 | These are best-effort numbers written by the ingest skill at create/update time. `wiki-lint` recomputes them and flags drift. The block is optional — pages without it are treated as fully extracted by convention. |
| 348 | |
| 349 | ## Typed Relationships |
| 350 | |
| 351 | Plain `[[wikilinks]]` in page bodies carry no semantic weight — they indicate "related to" but not *how*. The optional `relationships:` frontmatter block adds typed, directional edges to the knowledge graph. |
| 352 | |
| 353 | ### The `relationships:` block |
| 354 | |
| 355 | |
| 356 | relationships: |
| 357 | - target: "[[Transformer Architecture]]" |
| 358 | type: extends |
| 359 | - target: "[[LSTM]]" |
| 360 | type: contradicts |
| 361 | - target: "[[Attention Mechanism]]" |
| 362 | type: implements |
| 363 | |
| 364 | |
| 365 | Each entry has two required fields: |
| 366 | `target` — a wikilink (using the same format as `OBSIDIAN_LINK_FORMAT`) to the related page |
| 367 | `type` — one of the allowed semantic types below |
| 368 | |
| 369 | ### Allowed relationship types |
| 370 | |
| 371 | The table below is the framework default allowlist. A vault's `AGENTS.md` may extend it; consumers must use the effective allowlist and preserve owner semantics without coercion. |
| 372 | |
| 373 | | Type | Meaning | Example | |
| 374 | |---|---|---| |
| 375 | | `extends` | This page builds on or generalises the target | GPT extends Transformer Architecture | |
| 376 | | `implements` | This page is a concrete realisation of the target concept | BERT implements Masked Language Modelling | |
| 377 | | `contradicts` | This page's claims conflict with or refute the target | Evidence A contradicts Evidence B | |
| 378 | | `derived_from` | This page is based on or adapted from the target | Fine-tuning is derived from Transfer Learning | |
| 379 | | `uses` | This page depends on or relies on the target | RAG uses Vector Databases | |
| 380 | | `replaces` | This page supersedes or deprecates the target | GPT-4 replaces GPT-3 | |
| 381 | | `related_to` | Catch-all: related but no stronger directional type applies | Concept A is related to Concept B | |
| 382 | |
| 383 | ### Rules |
| 384 | |
| 385 | **Optional field** — omit the block entirely if no typed relationships are known. Untagged wikilinks remain valid and are treated as `related_to` by `wiki-export`. |
| 386 | **Don't duplicate** — if `[[foo]]` already appears as an inline wikilink, the `relationships:` entry just enriches it with a type; it is not a second link. |
| 387 | **Direction matters** — the page declaring the entry is the *source*; `target` is the destination. Only declare relationships from this page's perspective. |
| 388 | **Don't fabricate** — only add a typed entry when the source material makes the relationship direction and type clear. When in doubt, use `related_to` or omit. |
| 389 | |
| 390 | Skills that read `relationships:`: `wiki-export` (emits typed edges), `cross-linker` (writes typed entries when inferring links), `wiki-query` (surfaces type in answers and walks the typed-edge graph for multi-hop "how is X connected to Y" path queries — bounded BFS over the `relationships:` adjacency, frontmatter-only). |
| 391 | |
| 392 | ## Confidence and Lifecycle |
| 393 | |
| 394 | Every page carries two orthogonal trust signals plus an optional supersession link. |
| 395 | |
| 396 | The requiredness and lifecycle values below are framework defaults. A vault's `AGENTS.md` may extend lifecycle values or make trust fields optional. Validators must apply that effective owner schema while still validating any trust value that is present. |
| 397 | |
| 398 | The deterministic lint/trust consumer accepts owner schema through `OBSIDIAN_ALLOWED_LIFECYCLES`, `OBSIDIAN_ALLOWED_RELATIONSHIP_TYPES`, `OBSIDIAN_REQUIRED_TRUST_FIELDS`, and `OBSIDIAN_SCHEMA_SOURCE`. Resolution precedence is CLI > environment/config > these framework defaults (with lifecycle and relationship extensions additive). Explicit blank or whitespace-only values fail closed; omit the variable to select defaults. `wiki-lint/SKILL.md` owns the operational invocation contract. |
| 399 | |
| 400 | ### Required fields |
| 401 | |
| 402 | |
| 403 | base_confidence: 0.65 # [0.0, 1.0] — time-independent quality estimate. Stored once, recomputed on content change. |
| 404 | lifecycle: draft # draft | reviewed | verified | disputed | archived |
| 405 | lifecycle_changed: 2024-03-15 # ISO date of last state transition |
| 406 | # lifecycle_reason: "..." # optional free-text — why the state changed; surfaced by wiki-query |
| 407 | # superseded_by: "[[new-page]]" # wikilink; only when lifecycle=archived |
| 408 | |
| 409 | |
| 410 | `lifecycle_reason` and `superseded_by` are optional. Never fabricate them. |
| 411 | |
| 412 | ### Confidence formula |
| 413 | |
| 414 | The formula is a **manual base score**, not a deterministic URL classifier: |
| 415 | |
| 416 | |
| 417 | base_confidence = lineage_count_score * 0.5 + source_quality_score * 0.5 |
| 418 | |
| 419 | lineage_count_score = min(independent_evidence_lineages / 3, 1.0) |
| 420 | source_quality_score = avg(reviewed quality score per independent lineage) |
| 421 | |
| 422 | |
| 423 | After calculating the raw score, assess whether the evidence covers the page's material claims. Partial coverage may justify keeping or lowering the score; unsupported material claims require source/claim repair before any confidence change. Avoid small score churn without meaningful epistemic change. |
| 424 | |
| 425 | **Source-quality scores** (use the highest-matching bucket): |
| 426 | |
| 427 | | Bucket | Score | Examples | |
| 428 | |---|---|---| |
| 429 | | `paper` | 1.0 | arXiv, conference proceedings | |
| 430 | | `official` | 0.9 | `*.gov`, vendor docs | |
| 431 | | `documentation` | 0.85 | well-maintained third-party docs | |
| 432 | | `book` | 0.8 | books, technical references | |
| 433 | | `repository` | 0.75 | content-addressed repository/code evidence | |
| 434 | | `blog` | 0.55 | personal blogs | |
| 435 | | `session_transcript` | 0.5 | conversation history or completed operation | |
| 436 | | `forum` | 0.4 | Stack Overflow, HN, Reddit, issue-grade reports | |
| 437 | | `unknown` | 0.4 | catch-all/current config | |
| 438 | | `llm_generated` | 0.3 | LLM synthesis or unvalidated memory seed | |
| 439 | |
| 440 | **An independent evidence lineage** is an origin that can corroborate a claim independently. Canonical source IDs remain useful for identity, but identity alone does not prove independence. Collapse dependent evidence before counting: |
| 441 | |
| 442 | files, releases, and commits from one repository → one repository lineage; |
| 443 | retry/review/fix tasks in one workstream → one task lineage; |
| 444 | parent/child Kanban records → one task lineage; |
| 445 | byte-identical memories across profiles → one memory lineage; |
| 446 | a snapshot plus the mutable source it captures → one lineage; |
| 447 | aliases or metadata references resolving to one origin → one lineage. |
| 448 | |
| 449 | The deterministic `wiki-lint` path validates `_meta/trust-ledger.json`; it does not recompute confidence from source strings. New or materially changed pages are marked for manual review. Refresh the ledger only after explicit human approval. |
| 450 | |
| 451 | **Per-skill defaults** (ingest skills compute this automatically): |
| 452 | |
| 453 | | Skill | base_confidence | lifecycle | |
| 454 | |---|---|---| |
| 455 | | `wiki-ingest` (URL) | `0.17 + 0.5 × classify(url)` | `draft` | |
| 456 | | `wiki-ingest` (single doc) | per-source classifier | `draft` | |
| 457 | | `wiki-ingest` (multi-doc) | `min(N/3,1)×0.5 + avg_q×0.5` | `draft` | |
| 458 | | `wiki-research` | varies, often 0.85+ | `draft` | |
| 459 | | `wiki-capture` | 0.42 | `draft` | |
| 460 | | `*-history-ingest` | 0.42 | `draft` | |
| 461 | | `wiki-update` | 0.59 | `draft` | |
| 462 | | `wiki-synthesize` | `min(input_pages.base_confidence)` | `draft` | |
| 463 | |
| 464 | ### Lifecycle state machine |
| 465 | |
| 466 | Five states. **`stale` is not a state** — it is a computed overlay: `is_stale = (today − updated) > 90 days`. |
| 467 | |
| 468 | | State | Entered by | Notes | |
| 469 | |---|---|---| |
| 470 | | `draft` | Any ingest skill on first write | Default for all new pages | |
| 471 | | `reviewed` | Human edit only | | |
| 472 | | `verified` | Human edit only | Time alone never demotes verified pages | |
| 473 | | `disputed` | Manual edit only | Overrides every state except `archived` in display | |
| 474 | | `archived` | Manual edit, or ingest skill setting `superseded_by` | Terminal | |
| 475 | |
| 476 | Only ingest skills set `draft`. All other transitions require a human editor. Update `lifecycle_changed` whenever the state changes. |
| 477 | |
| 478 | Two edge classes are therefore **illegal** and are reported by `obsidian-wiki lint` as `illegal_lifecycle_transitions`: anything falling back to `draft` (`reviewed|verified|disputed → draft`), and any exit from `archived` (it is terminal — restoring a page is a deliberate delete-and-recreate, not a transition). The check compares against the lifecycle recorded in `_meta/trust-ledger.json` at the page's last review, so it only sees pages that have been reviewed at least once. |
| 479 | |
| 480 | ## Importance Tiering |
| 481 | |
| 482 | The `tier:` field controls which pages get updated on each ingest pass and their priority in retrieval. As wikis grow, re-reading every page on every ingest wastes tokens — tiering lets ingest and query skills focus effort where it matters most. |
| 483 | |
| 484 | ### Three tiers |
| 485 | |
| 486 | | Tier | Meaning | Ingest behavior | Query priority | |
| 487 | |---|---|---|---| |
| 488 | | `core` | Load-bearing pages — many other pages depend on them (high incoming-link count or bridge position). Always worth updating. | Always update if the source is even marginally relevant | Surfaced first in index and full-read passes | |
| 489 | | `supporting` *(default)* | Standard wiki pages with moderate connectivity | Update when the source has clear new claims for this page | Standard priority | |
| 490 | | `peripheral` | Low-connectivity pages — rarely linked, narrowly scoped | Skip unless the source is *primarily* about this topic | Last resort; skipped when trimming to context budget | |
| 491 | |
| 492 | ### Assignment rules |
| 493 | |
| 494 | **New pages:** default to `tier: supporting` |
| 495 | **Promote to `core`:** when a page accumulates ≥5 incoming wikilinks **or** is flagged as a bridge by `wiki-status` insights mode |
| 496 | **Demote to `peripheral`:** when a page has ≤1 incoming link and hasn't been updated in 90+ days |
| 497 | **Human override always wins** — edit `tier:` manually to lock a page at any level |
| 498 | Existing pages without `tier:` are treated as `supporting` (backward compatible — no migration needed) |
| 499 | |
| 500 | ### Who manages tier |
| 501 | |
| 502 | `wiki-ingest` reads `tier:` to decide whether to update a page on the current pass |
| 503 | `wiki-query` uses `tier:` to order candidates in the index pass and trim to context budget |
| 504 | `wiki-status` insights mode computes graph metrics and **suggests** tier assignments — it never writes them automatically |
| 505 | `wiki-lint` flags missing `tier:` on newly created pages (Phase 2 enforcement, same timeline as `base_confidence`) |
| 506 | |
| 507 | ## Retrieval Primitives |
| 508 | |
| 509 | Reading the vault is the dominant cost of every read-side skill. Use the cheapest primitive that can answer the question and **escalate only when the cheaper one is insufficient**. Any skill that needs content from the vault should follow this table rather than jumping straight to full-page reads. |
| 510 | |
| 511 | | Need | Primitive | Relative cost | |
| 512 | |---|---|---| |
| 513 | | Does a page exist? What's its title/category/tags? | Read `index.md`; `Grep` frontmatter blocks (scope with a pattern that targets `^---` blocks at file heads) | **Cheapest** | |
| 514 | | 1–2 sentence preview of a page | Read the `summary:` field in its frontmatter | **Cheap** | |
| 515 | | A specific claim or section inside a page | `Grep -A <n> -B <n> "<term>" <file>` — returns only the matching lines plus context | **Medium** | |
| 516 | | Whole-page content | `Read <file>` | **Expensive** — last resort | |
| 517 | | Relationships across pages | `Grep "\[\[.*?\]\]"` across the vault, or walk wikilinks from a known page | Case-by-case | |
| 518 | |
| 519 | **Search command preference:** for shell/file searches, use ripgrep (`rg`, `rg --files`) when available; if not, fall back to `grep`/`find`. Capitalized `Grep`/`Glob` names in these skills are tool-generic primitives for agents that expose those tools. |
| 520 | |
| 521 | **The rule:** escalate only when the cheaper primitive can't answer the question. If you can answer from `summary:` fields alone, don't read page bodies. If a grepped section with `-A 10 -B 2` gives you the claim, don't read the whole page. A 500-line page opened to read 15 lines is 485 lines of wasted tokens. |
| 522 | |
| 523 | **Why this matters:** a 20-page vault lets you get away with full-vault scans. A 200-page vault does not. The primitives above are how the skills framework scales to large vaults without a database. |
| 524 | |
| 525 | Skills that consume this table: `wiki-query`, `cross-linker`, `wiki-lint`, `wiki-status` (insights mode). Any new skill that reads the vault should cite this section rather than reinvent the pattern. |
| 526 | |
| 527 | ## QMD Index Freshness |
| 528 | |
| 529 | QMD is an optional search index layered on top of the vault. The markdown vault is the source of truth. Any skill that writes wiki markdown should refresh QMD after the vault write completes, but only when `QMD_WIKI_COLLECTION` is configured and the local QMD transport is available. If QMD refresh fails, keep the vault changes and report the QMD status separately. |
| 530 | |
| 531 | Use the cheapest verification path that proves the new content is visible: `qmd update`, `qmd embed` only if vectors are stale or missing, then a targeted `qmd get` or `qmd ls` check for one written page or the collection root. Read-only skills should not refresh QMD. |
| 532 | |
| 533 | ## Core Principles |
| 534 | |
| 535 | **Compile, don't retrieve.** The wiki is pre-compiled knowledge. When you ingest a source, update every relevant page — don't just create a summary of the source. |
| 536 | |
| 537 | **Compound over time.** Each ingest should make the wiki smarter, not just bigger. Merge new information into existing pages, resolve contradictions, strengthen cross-references. |
| 538 | |
| 539 | **Provenance matters.** Every claim should trace to a source. When updating a page, note which source prompted the update. |
| 540 | |
| 541 | **Mark inferences.** Default sentences are extracted. Mark synthesized claims with `^[inferred]` and contested claims with `^[ambiguous]`. A wiki that hides its guessing rots silently; one that marks it stays trustworthy. |
| 542 | |
| 543 | **Human curates, LLM maintains.** The human decides what sources to add and what questions to ask. The LLM handles the bookkeeping — updating cross-references, maintaining consistency, noting contradictions. |
| 544 | |
| 545 | **Obsidian is the IDE.** The user browses and explores the wiki in Obsidian. Everything must be valid Obsidian markdown with working wikilinks. |
| 546 | |
| 547 | ## Link Format |
| 548 | |
| 549 | All internal links connecting wiki pages are controlled by `OBSIDIAN_LINK_FORMAT` from the resolved config (default: `wikilink`). |
| 550 | |
| 551 | | Setting | Syntax | Example | |
| 552 | |---|---|---| |
| 553 | | `wikilink` *(default)* | `[[path/to/page]]` or `[[path/to/page\|display text]]` | `[[concepts/foo\|foo]]` | |
| 554 | | `markdown` | `[display text]` | `[foo]` | |
| 555 | |
| 556 | ### Generating markdown-format links |
| 557 | |
| 558 | When `OBSIDIAN_LINK_FORMAT=markdown`: |
| 559 | Compute the path from the **current file's directory** to the **target `.md` file** using `..` to climb up as needed. |
| 560 | Use the page title or a natural phrase as display text. |
| 561 | Always include the `.md` extension. |
| 562 | |
| 563 | | Current file | Target | Relative link | |
| 564 | |---|---|---| |
| 565 | | `index.md` | `concepts/foo.md` | `[foo]` | |
| 566 | | `concepts/foo.md` | `entities/bar.md` | `[bar]` | |
| 567 | | `projects/my-project/my-project.md` | `concepts/foo.md` | `[foo]` | |
| 568 | | `projects/my-project/concepts/arch.md` | `entities/bar.md` | `[bar]` | |
| 569 | |
| 570 | The `[[path\|display text]]` wikilink form maps to `[display text]` in Markdown mode. |
| 571 | |
| 572 | **Scope:** this setting affects only newly written or updated links. Existing vault content is never automatically migrated — users who want to convert old links can run the `cross-linker` or `wiki-lint` skill. |
| 573 | |
| 574 | Every write skill reads `OBSIDIAN_LINK_FORMAT` from config before generating links and applies the correct format. |
| 575 | |
| 576 | ## Config Resolution Protocol |
| 577 | |
| 578 | **All skills must resolve config using this algorithm — do not hard-code `.env` or the global config path directly.** This ensures single-vault, multi-vault, project-local, and VPS setups all work correctly. |
| 579 | |
| 580 | ### Global config directory |
| 581 | |
| 582 | The global config directory is **XDG-style**: `$XDG_CONFIG_HOME/obsidian-wiki` (default `~/.config/obsidian-wiki`). Installs that already have a `~/.obsidian-wiki` directory keep using it — so an existing setup never breaks — but any **new** install lands under the XDG path. Resolve it with: |
| 583 | |
| 584 | |
| 585 | obsidian_wiki_config_dir() { |
| 586 | local xdg_dir="${XDG_CONFIG_HOME:-$HOME/.config}/obsidian-wiki" |
| 587 | local legacy_dir="$HOME/.obsidian-wiki" |
| 588 | if [[ -d "$legacy_dir" && ! -e "$xdg_dir" ]]; then |
| 589 | echo "$legacy_dir" |
| 590 | else |
| 591 | echo "$xdg_dir" |
| 592 | fi |
| 593 | } |
| 594 | |
| 595 | |
| 596 | Everywhere below, "the global config dir" means `$(obsidian_wiki_config_dir)`, and "the global config" means `$(obsidian_wiki_config_dir)/config`. |
| 597 | |
| 598 | ### Resolution order |
| 599 | |
| 600 | **Inline vault override (`@name`)** — if the user's request contains an `@<name>` token (e.g. `@work save this`, `query @personal about X`), resolve `<global config dir>/config.<name>` directly and use its `OBSIDIAN_VAULT_PATH`. This **overrides** both the CWD `.env` walk-up and the active symlink, and applies to **that invocation only** — never run `ln -sf` or otherwise change the active vault for an `@name` request. If `<global config dir>/config.<name>` doesn't exist, tell the user it doesn't exist and list the available vaults (the `wiki-switch` **List** logic), then stop — do **not** silently fall back to the default. The `@name` is a routing directive, not content: strip it out before treating the rest of the request as the actual instruction or page text. |
| 601 | **Walk up from CWD** — look for a `.env` file in the current directory, then each parent, up to `$HOME`. Stop at the first `.env` that contains `OBSIDIAN_VAULT_PATH`. If its value is empty, stop there too: tell the user which `.env` blocked resolution (a blank line copied from `.env.example` does this) instead of falling through to the global config. |
| 602 | **Global config** — if no local `.env` found, read the global config (`$(obsidian_wiki_config_dir)/config`). |
| 603 | **Prompt setup** — if neither exists, tell the user: "No config found. Run `wiki-setup` to initialize your wiki." |
| 604 | |
| 605 | `@name` is a **per-invocation override** — it targets one vault for one request. `/wiki-switch <name>` is the **persistent default** — it re-points the active symlink for all future requests. Use `@name` to touch the other vault from anywhere without disturbing your default ("brain") vault. |
| 606 | |
| 607 | |
| 608 | find_config() { |
| 609 | # $1 = parsed @name from the request, if any (else empty) |
| 610 | local config_dir |
| 611 | config_dir="$(obsidian_wiki_config_dir)" |
| 612 | if [[ -n "$1" ]]; then |
| 613 | [[ -f "$config_dir/config.$1" ]] && { echo "$config_dir/config.$1"; return; } |
| 614 | echo ""; return # named vault missing → caller reports + lists, no fallback |
| 615 | fi |
| 616 | dir="$PWD" |
| 617 | while [[ "$dir" != "$HOME" && "$dir" != "/" ]]; do |
| 618 | [[ -f "$dir/.env" ]] && grep -q "OBSIDIAN_VAULT_PATH" "$dir/.env" && { echo "$dir/.env"; return; } |
| 619 | dir="$(dirname "$dir")" |
| 620 | done |
| 621 | [[ -f "$config_dir/config" ]] && { echo "$config_dir/config"; return; } |
| 622 | echo "" |
| 623 | } |
| 624 | |
| 625 | |
| 626 | ### Vault-scoped state |
| 627 | |
| 628 | Skills that write runtime state (e.g. `daily-update`) must scope that state to the resolved vault, not to a global path. Use: |
| 629 | |
| 630 | |
| 631 | VAULT_ID=$(echo "$OBSIDIAN_VAULT_PATH" | md5sum 2>/dev/null || md5 -q - <<< "$OBSIDIAN_VAULT_PATH" | cut -c1-8) |
| 632 | STATE_DIR="$(obsidian_wiki_config_dir)/state/$VAULT_ID" |
| 633 | |
| 634 | |
| 635 | ### Standard "Before You Start" block |
| 636 | |
| 637 | Every skill's setup section should read: |
| 638 | |
| 639 | > **Resolve config** — follow the Config Resolution Protocol in `llm-wiki/SKILL.md`. Honor an inline `@name` override first, then walk up from CWD for `.env`, fall back to the global config, else prompt setup. This gives `OBSIDIAN_VAULT_PATH` and any tool-specific path overrides. |
| 640 | |
| 641 | ## Writing Profile Resolution |
| 642 | |
| 643 | Before drafting or rewriting natural-language Markdown, resolve the global config directory with the XDG/legacy algorithm above, then read `<global config dir>/WRITING.md` when it exists. A missing or empty `WRITING.md` means there are no custom writing preferences. If that optional read fails, warn and continue with the default framework guidance. |
| 644 | |
| 645 | The effective precedence is framework invariants > current task/skill requirements > current project `AGENTS.md` > vault `AGENTS.md` > global `WRITING.md`. Framework invariants include schema, provenance, and safety; operation-specific requirements remain authoritative for the current task. Unspecified project and vault rules are inherited from less-specific layers, and more specific same-topic rules win. |
| 646 | |
| 647 | Writing preferences apply only to newly drafted or rewritten natural-language fields and body content. This includes natural-language title and summary values in YAML frontmatter, but preferences cannot alter YAML syntax, required keys, structure, types, or machine-generated fields. JSON, structured logs, and pass-through content remain unchanged and retain their required formats and source fidelity. |
| 648 | |
| 649 | ## Environment Variables |
| 650 | |
| 651 | The wiki is configured through environment variables (see `.env.example`). The only required variable is the vault path — everything else has sensible defaults. |
| 652 | |
| 653 | `OBSIDIAN_VAULT_PATH` — Where the wiki lives **(required)** |
| 654 | `OBSIDIAN_SOURCES_DIR` — Where raw source documents are |
| 655 | `OBSIDIAN_CATEGORIES` — Comma-separated list of categories |
| 656 | `WIKI_SKIP_PROJECTS` — Comma-separated substrings; any project dir whose name contains one is excluded from history ingest (scan + delta + manifest). See the "Project Scoping" step in the history-ingest skills. |
| 657 | `CLAUDE_HISTORY_PATH` — Where to find Claude conversation data |
| 658 | `CODEX_HISTORY_PATH` — Where to find Codex session data |
| 659 | `HERMES_HOME` — Where to find Hermes agent data |
| 660 | `OPENCLAW_HOME` — Where to find OpenClaw data |
| 661 | `COPILOT_HISTORY_PATH` — Where to find Copilot session data |
| 662 | `OBSIDIAN_LINK_FORMAT` — Internal link syntax: `wikilink` (default) or `markdown` |
| 663 | `WIKI_TOKEN_WARN_THRESHOLD` — Emit a warning in `wiki-status` when the full-wiki token estimate exceeds this value (default: `100000`). Set to `0` to disable. See `wiki-status` for the token footprint report. |
| 664 | `WIKI_STAGED_WRITES` — When `true`, all LLM-written pages go to `_staging/<category>/` for human review before promotion. See `wiki-setup` and `wiki-stage-commit` for details. |
| 665 | `CODE_UNDERSTANDING_BACKEND` — how wiki-update understands a project before distilling: `auto` (CodeGraph when available, else builtin ast-extract + rg; default), `builtin`, or `codegraph` (explicitly require; warn/error if unavailable). |
| 666 | `CODE_UNDERSTANDING_CODEGRAPH_BIN` — optional path to the codegraph binary when it isn't on PATH. |
| 667 | `CODE_UNDERSTANDING_CODEGRAPH_BIN` — optional path to the codegraph binary when it isn't on PATH. |
| 668 | Both resolve like `OBSIDIAN_VAULT_PATH`: a real environment variable wins (empty counts as |
| 669 | unset), then the nearest `.env` walking up from the project directory, then the global config |
| 670 | (`$(obsidian_wiki_config_dir)/config`), then the default. |
| 671 | `OBSIDIAN_MAX_PAGES_PER_INGEST` — cap on pages created/updated per `wiki-ingest` run (default: `15`). See `wiki-ingest`, Step 4. |
| 672 | `LINT_SCHEDULE` — how often `daily-update` also runs `wiki-lint`: `daily` \| `weekly` (default) \| `manual`. See `daily-update`, Step 4a. |
| 673 | |
| 674 | No API keys are needed — the agent running these skills already has LLM access built in. |
| 675 | |
| 676 | ## Modes of Operation |
| 677 | |
| 678 | The wiki supports three ingest modes: |
| 679 | |
| 680 | | Mode | When to use | What happens | |
| 681 | |---|---|---| |
| 682 | | **Append** | Small delta, incremental updates | Compute delta via manifest, ingest only new/modified sources | |
| 683 | | **Rebuild** | Major drift, fresh start needed | Archive current wiki to `_archives/`, clear, reprocess all sources | |
| 684 | | **Restore** | Need to go back | Bring back a previous archive | |
| 685 | |
| 686 | Use `wiki-status` to see the delta and get a recommendation. Use `wiki-rebuild` for archive/rebuild/restore operations. |
| 687 | |
| 688 | ## Reference |
| 689 | |
| 690 | For details on specific operations, see the companion skills: |
| 691 | **wiki-status** — Audit what's ingested, compute delta, recommend append vs rebuild |
| 692 | **wiki-rebuild** — Archive current wiki, rebuild from scratch, or restore from archive |
| 693 | **wiki-ingest** — Distill source documents into wiki pages and raw text/chat/log data |
| 694 | **claude-history-ingest** — Ingest Claude conversation history |
| 695 | **codex-history-ingest** — Ingest Codex CLI session history |
| 696 | **wiki-query** — Answer questions against the wiki |
| 697 | **wiki-lint** — Audit and maintain wiki health |
| 698 | **wiki-setup** — Initialize a new vault |
| 699 |
Discussion
Alternatives
Browse more free Claude skills or everything in Development.