Miru code agent

Code search that lets AI coding agents find the right files in fewer steps.

by takara-ai·MIT license·★ 11 Stars on the repo·GitHub ↗

Files of Miru code

takara-ai/main1 file
README.md
Show the full text457 lines

Miru (見る)

CI coverage license bun

Hybrid code search for AI coding agents. Find code by meaning, not grep.

Your AI agent finds the code it needs with up to 50% fewer tokens.

Miru returns the best chunks (path, lines, snippet) for questions like "where is auth middleware configured?" — plugged directly into Claude Code, Cursor, Copilot, Codex, and 9+ other agents via MCP.


How it fits into your agent's workflow

How Miru fits into an agent's workflow

Miru replaces the grep/glob-style search agents fall back on today. Install once, and every connected agent gets search and find_related MCP tools automatically.

Install via a plugin marketplace

Prefer this when your IDE has a plugin marketplace — no bun add -g step, and updates go through the IDE's own plugin flow. Run miru setup first if you have not authenticated yet (see Set up credentials).

Claude Code:

/plugin marketplace add takara-ai/miru-code
/plugin install miru

Restart Claude Code (or reload plugins) when prompted. Update: /plugin marketplace update miru then reinstall. Remove: /plugin uninstall miru.

Cursor:

  1. Dashboard → Plugins → Team Marketplaces → Add Marketplace → Import from Repo → takara-ai/miru-code
  2. Customize (sidebar) → find miru → Install → choose project or user scope

Any other IDE, or if marketplace install isn't available: use the CLI install below.

What you get from a plugin install vs the CLI

Plugin installs don't carry full CLI parity — what each one gives you depends on the IDE's own plugin capabilities, not just Miru's packaging:

Claude Code Codex Cursor
MCP tools (search, locate, expand, find_related) ✅ ✅ ❌ (plugin ships skills + rules only — no MCP entry yet)
miru / Caveman / STE skills ✅ ✅ ✅
Dedicated sub-agent (miru:miru-code) ✅ — —
Benchmark mode toggle (/plugin configure) ✅ — —
Credentials own plugin-scoped dir own plugin-scoped dir n/a

— means the IDE's plugin schema has no equivalent mechanism to port these to (not a packaging gap we can close): Codex's and Cursor's plugin manifests have no agents or userConfig fields, so the dedicated sub-agent and the benchmark toggle are Claude-Code-only.

Credentials are plugin-scoped, not shared with your CLI install. Claude Code and Codex plugins each store auth in their own IDE-managed data directory (survives plugin updates, removed on uninstall). This means miru setup done via the CLI does not carry over to a plugin install, or vice versa — each authenticates independently on first use (interactive device-code login bootstraps automatically). If you use both the CLI and a plugin, you'll sign in twice.

Install (CLI)

bun add -g @takara-ai/miru-code

Set up credentials

miru setup

Interactive miru setup defaults to device-code login and saves the resulting credentials locally. Manual bearer-token entry is still available with --key. If credentials are missing, the interactive MCP/plugin path can bootstrap the same device flow automatically on first use.

miru setup --device           # explicit device-code login
miru setup --key YOUR_TOKEN   # store a bearer token directly
miru setup --clear            # remove stored credentials

Miru stores versioned credentials in credentials.json and automatically loads or refreshes them for MCP and CLI use. TAKARA_API_KEY still overrides stored credentials when set explicitly.

Add to your IDE

miru install

Interactive TUI — ↑↓ move, space toggle, a all, enter confirm. Pick agents and integrations:

Integration What it does
MCP server search, locate, expand, and find_related tools in the agent
Instructions Search policy in CLAUDE.md / AGENTS.md / GEMINI.md
Sub-agent Dedicated miru-code agent file
Cursor rules Always-on .cursor/rules/miru-code.mdc (Cursor only)
Old Miru hooks Removes stale search hooks from older releases
Caveman (experimental) On-demand chat compression skill (/caveman)
STE writing (experimental) On-demand clear technical English for docs (/ste)

Restart the IDE when done.

miru uninstall   # remove miru config

Supported: Cursor · Claude Code · Gemini CLI · Kiro · OpenCode · GitHub Copilot · Codex · VS Code · Visual Studio (Windows) · Windsurf / Devin Desktop

IDE MCP Instructions / rules Caveman (experimental) STE (experimental)
Cursor ~/.cursor/mcp.json ~/.cursor/rules/miru-code.mdc ~/.agents/skills/caveman/SKILL.md ~/.agents/skills/ste/SKILL.md
Claude Code ~/.claude.json ~/.claude/CLAUDE.md ~/.claude/skills/caveman/SKILL.md ~/.claude/skills/ste/SKILL.md
Gemini CLI ~/.gemini/settings.json ~/.gemini/GEMINI.md ~/.agents/skills/caveman/SKILL.md ~/.agents/skills/ste/SKILL.md
Kiro ~/.kiro/settings/mcp.json ~/.kiro/steering/miru.md ~/.kiro/skills/caveman/SKILL.md ~/.kiro/skills/ste/SKILL.md
OpenCode $XDG_CONFIG_HOME/opencode/opencode.json(c) (else ~/.config/opencode/…) …/AGENTS.md ~/.agents/skills/caveman/SKILL.md ~/.agents/skills/ste/SKILL.md
GitHub Copilot ~/.copilot/mcp-config.json — ~/.agents/skills/caveman/SKILL.md ~/.agents/skills/ste/SKILL.md
Codex ~/.codex/config.toml ~/.codex/AGENTS.md ~/.agents/skills/caveman/SKILL.md ~/.agents/skills/ste/SKILL.md
VS Code …/Code/User/mcp.json — ~/.agents/skills/caveman/SKILL.md ~/.agents/skills/ste/SKILL.md
Visual Studio %USERPROFILE%\.mcp.json — ~/.agents/skills/caveman/SKILL.md ~/.agents/skills/ste/SKILL.md
Windsurf — — ~/.agents/skills/caveman/SKILL.md ~/.agents/skills/ste/SKILL.md
Plugin packaging
  • Codex: .codex-plugin/plugin.json, .agents/plugins/marketplace.json, shares the root mcp.json
  • Claude Code: .claude-plugin/plugin.json, .claude-plugin/marketplace.json, its own .claude-plugin/mcp.json (needed for the userConfig benchmark toggle — see What you get from a plugin install vs the CLI)
  • Cursor: plugin.json and .cursor/rules/miru-code-search.mdc

Current limitation:

  • these plugin manifests launch the published Miru runtime through bunx @takara-ai/miru-code@latest mcp
  • that means local source edits do not affect plugin behavior until a package version is published
  • and a fully self-contained “no Bun required” plugin install is still future work
On-demand skills

Sub-agent files are also written where supported (see miru install plan). Windsurf hooks only (experimental) — no MCP entry yet. Caveman is an on-demand Agent Skill (default off): invoke with /caveman or “talk like caveman”; stop with “normal mode”. Invocation UI varies by IDE (/caveman, $caveman, @caveman, etc.). Most IDEs share ~/.agents/skills/caveman/SKILL.md (including Copilot / VS Code / Visual Studio); Claude Code and Kiro keep vendor-native skill dirs. Ownership is tracked on the shared path so uninstalling one IDE keeps the skill while another still owns it; selecting all owners removes it once. STE is an on-demand Agent Skill (default off): invoke with /ste or “de-slop this”; keep articles and complete sentences. Most IDEs share ~/.agents/skills/ste/ the same way (ownership via miru-owners.json); Claude Code and Kiro keep vendor-native STE dirs.

STE writing (experimental)

STE helps write clear technical English for docs, runbooks, errors, and release notes (pragmatic ASD-STE100-inspired rules). Invoke with /ste or “de-slop this”. Keep articles and complete sentences — not telegraph-style omission.

Not ASD-certified. Full dictionary compliance needs the official standard at asd-ste100.org. Miru does not ship the copyrighted ASD dictionary.

Not for marketing or brand copy. Off by default; enable STE in the installer for any supported IDE. Restart the IDE (or reload skills) after install. For Codex, install also sets [features] skills = true in ~/.codex/config.toml (same as Caveman).

Caveman mode (experimental)

Caveman compresses live chat replies (less filler, max meaning). Intensities: /caveman lite|full|ultra (default full). Persisted artifacts (commits, PRs, customer docs) stay normal prose unless you ask otherwise.

Security / destructive warnings use clear normal prose (auto-clarity) — brevity never hides risk. Session token savings vary; the skill itself costs input tokens. No guaranteed %.

Off by default at install time. Enable Caveman in the installer for any supported IDE. Restart the IDE (or reload skills) after install. For Codex, the installer also sets [features] skills = true in ~/.codex/config.toml (required for Codex to load skills).

Team sub-agent in a repo (optional):

miru init --agent claude --force

Try it

Meaning-based questions → search. Exact strings (env vars, symbols, error codes) → locate.

miru search "auth middleware" ./src
miru locate REDIS_HOST ./src
miru expand src/auth.ts 42 ./src
miru find-related src/auth.ts 42 ./src

Terminal output is human-readable; use --json for scripts. One-off without installing:

bunx @takara-ai/miru-code@latest search "auth middleware" ./src

MCP tools

When wired via miru install, the MCP server exposes search, locate, expand, and find_related. read_benchmark appears only in benchmark mode. repo is optional: it defaults to the server's startup directory. Set it for another repo, or if your MCP client starts elsewhere. The index is built on the first call and cached for the session.

Tool When to use
search Default for code exploration — hybrid semantic + keyword search. One call per question.
locate Exact substrings (env vars, symbols, error codes) — prefer over Grep.
expand More context in the same file when a hit has truncated: true.
find_related Similar code in other files from a file_path + anchor_line.
read_benchmark Cumulative token-savings rollup (benchmark mode only).
Workflow
  1. search with query — returns compact snippets (~±15 lines) and relevance scores.
  2. If a hit has truncated: true, call expand with file_path and anchor_line — not another search or a full-file read.
  3. To trace similar patterns elsewhere, call find_related with the same file_path and anchor_line.
  4. Use your editor's Read on absolute_path only when editing or when expand still lacks context.

Keep repo on follow-up calls when set.

Prefer these tools over Grep, Glob, or SemanticSearch when Miru MCP is connected — hooks and instructions enforce that when enabled.

Local repo hits include absolute_path for one-click navigation. Parameter reference is below under MCP parameters.

How it works

Hybrid search: Takara embeddings + BM25 + fusion + reranking. Index code, docs, config, or all with --content.

MCP watches local files and updates the index incrementally. Package upgrades invalidate stale caches via the version epoch.

OS Index cache
macOS ~/Library/Caches/miru
Linux ~/.cache/miru
Windows %LOCALAPPDATA%\miru\Cache
Chunking & languages

Miru chunks source in tiers: AST (tree-sitter, default) → structural heuristics → line splits.

AST chunking — 26 languages (syntax-aware boundaries via vendored web-tree-sitter grammars):

Language Typical extensions
astro .astro
bash .sh, .bash, .zsh
c .c
cpp .cpp, .h, .hpp, etc.
csharp .cs
css .css
dart .dart
elixir .ex, .exs
embeddedtemplate .erb, .ejs
go .go
haskell .hs
html .html, .htm
java .java
javascript .js, .jsx, .mjs, .cjs
json .json
ocaml .ml, etc.
php .php
python .py, .pyi
ruby .rb
rust .rs
scala .scala
solidity .sol
sql .sql
svelte .svelte
typescript .ts, .tsx, .mts, .cts
vue .vue

Structural fallback (brace/indent heuristics when AST is unavailable): python, go, typescript, javascript, cpp, c.

Line fallback: everything else that gets indexed (kotlin, swift, etc.) — still searchable, coarser chunks.

Set MIRU_AST_CHUNKING=0 to disable AST and use structural → lines only.

CLI reference

Run miru in a terminal or miru -h for the command list, or miru <command> -h for details.

Command Purpose
miru setup Authenticate and store credentials
miru install Configure IDE (global)
miru uninstall Remove IDE config
miru search <query> [path] Search (-k N, --content, --json)
miru locate <literal> [path] Exact substring in the index
miru expand <file> <line> [path] Adjacent chunks in the same file
miru find-related <file> <line> [path] Related chunks
miru benchmark on/off/status/clear Toggle MCP benchmark mode / clear report
miru init --agent <id> Project-local sub-agent
miru clear [path] Drop index cache (use after big CLI-only refactors)
miru mcp Start MCP server (--benchmark for comparisons)

CLI uses hyphens (find-related); MCP tool names use underscores (find_related).

Library

bun add @takara-ai/miru-code
import { MiruIndex } from "@takara-ai/miru-code";

const index = await MiruIndex.fromPath("./src");
const results = await index.search({ query: "BM25 tokenize" });

Environment

Variable Notes
TAKARA_API_KEY Required
MIRU_OPENAI_BASE_URL Default https://infer.takara.ai/v1
MIRU_OPENAI_EMBEDDING_MODEL Default ds1-miru-int8
MIRU_WORKSPACE_ROOT Optional: restrict MCP local repo paths to this directory
MIRU_MAX_INDEX_FILES Cap files indexed per operation
MIRU_ALLOW_HTTP_GIT Set 1 to allow plain http:// git clones
MIRU_MCP_WATCH Set 0 to disable MCP file watch
MIRU_AST_CHUNKING Set 0 to disable tree-sitter AST chunking
MIRU_BENCHMARK_HISTORY_PATH Override; see Benchmark mode
MIRU_CACHE_HOME Override the platform cache root (see How it works)
MIRU_QUIET Set 1 to skip the framed CLI banner (subtitle only on color terminals)
NO_COLOR Disable CLI colors

See .env.example for more.

Privacy and API usage

Miru sends file contents to the Takara inference API when building an index and when embedding search queries. Chunks from your repo are transmitted over HTTPS to generate embeddings. API usage may incur cost depending on your Takara plan.

If you index proprietary code, make sure that sending snippets to Takara's endpoint fits your security and compliance requirements. MIRU_WORKSPACE_ROOT is an opt-in boundary for MCP local repo paths only, and restricts indexing to a single workspace directory when set.

Enterprise self-hosted embeddings (no Takara egress): see docs/self-hosted-sagemaker.md.

Benchmark mode

Optional measurement of how many tokens Miru saves versus a simple Grep workflow (ripgrep + reading the top matched file). Useful when evaluating Miru; leave it off day-to-day. Only local repo paths are compared — git URL repos skip the comparison and return benchmark_skipped: "local_repo_only".

miru benchmark on       # add --benchmark to installed MCP configs
miru benchmark status
miru benchmark off      # prefer off when finished measuring
miru benchmark clear    # delete the global report

Restart agents after changing mode. miru install keeps --benchmark if it was already enabled.

While on, each search / locate response includes a compact benchmark block (save_pct, miru_tok, grep_tok, saved_tok, rank1). Call read_benchmark for a cumulative rollup (or ask the agent when you want totals).

History is global (not per-repo) under Miru's state directory:

OS Default path
macOS ~/Library/Application Support/miru/benchmark-history.json
Linux ~/.config/miru/benchmark-history.json ($XDG_CONFIG_HOME/miru/… when set)
Windows %APPDATA%\miru\benchmark-history.json

Override with MIRU_BENCHMARK_HISTORY_PATH. Append-only JSONL of compact token deltas (no query text); read_benchmark returns cumulative totals. Stored in plaintext — miru benchmark clear or uninstall on shared machines.

MCP parameters

**search**

Param Required Notes
query yes Natural language or code query
repo no Startup directory by default; set for another repo
include no Gitignore-style glob patterns; only matching files are searched (same as locate.include)
exclude no Gitignore-style glob patterns; matching files are skipped (same as locate.exclude)
dedupe_by_file no Keep best hit per file (default true)

**locate**

Param Required Notes
literal yes Exact substring to find
repo no Startup directory by default; set for another repo
include no Gitignore-style glob patterns; only matching files are searched
exclude no Gitignore-style glob patterns; matching files are skipped
mode no count · locations · lines (default). Prefer count/locations when possible
limit no Cap returned hits. Omit to return all matches
ignore_case no Case-insensitive match (default false)

**expand**

Param Required Notes
file_path yes From hit file_path or absolute_path (local repos)
anchor_line yes From the search hit (anchor_line when truncated, else start_line)
repo no Same repo as the search, if set
before / after no Extra chunks before/after anchor (default 1 each)

**find_related**

Param Required Notes
file_path yes From a search hit
anchor_line yes From the search hit
repo no Same repo as the search, if set

**read_benchmark** (benchmark mode only)

Param Required Notes
repo no Filter rollup to one local path or git URL. Omit for all saved queries

Manual MCP (skip miru install)

{
  "miru": {
    "command": "miru",
    "args": ["mcp"]
  }
}

For benchmark mode, set "args": ["mcp", "--benchmark"] (or append the flag). Prefer miru benchmark on after a normal install — it updates every agent config.

Run miru setup once so the server can load credentials from credentials.json. If the MCP server starts in an interactive terminal without stored credentials, it will start device login automatically.

Use bunx + @takara-ai/miru-code@latest if miru is not global. The installer uses this command so each new MCP server launch can pick up a published version without a global update. A running server keeps its current version until restarted. Wrapper key varies by IDE (mcpServers, servers, or mcp).

Older headless MCP configs that launch miru without a subcommand, with or without MCP flags, continue to work. Run miru install again to update managed entries to the explicit mcp subcommand, then restart your coding agent. For manual MCP configs, add mcp after the executable or package name.

Developing

git clone https://github.com/takara-ai/miru-code.git && cd miru-code
bun install && cp .env.example .env.local
bun test && bun run typecheck

See CONTRIBUTING.md for pre-commit hooks, commit message conventions, and the PR process.

Local MCP: "command": "bun", "args": ["/path/to/miru-code/src/cli.ts"]

miru -h / setup print a framed wordmark on color terminals (MIRU_QUIET=1 for subtitle only). Crane art lives in src/brand-banner.ts; regenerate with bun run scripts/render-crane-art.ts (ImageMagick). The crane is a registered mark of Takara.ai Ltd.

Credits

Miru uses work by MinishLab. Thank you to its authors.

  • semble (MIT). Miru ports parts of semble. These parts include the file walker, the index layout, the BM25 tokenizer, the hybrid search pipeline, and the ranking signals.
  • potion-code and Model2Vec (MIT). The WordPiece tokenizer in tokenizer/tokenizer.json comes from potion-code.

See NOTICE for the licence text and the Model2Vec citation.

License

MIT

1# Miru (見る)
2 
3[![CI](https://github.com/takara-ai/miru-code/actions/workflows/ci.yml/badge.svg)](https://github.com/takara-ai/miru-code/actions/workflows/ci.yml) [![coverage](https://img.shields.io/endpoint?url=https%3A%2F%2Fraw.githubusercontent.com%2Ftakara-ai%2Fmiru-code%2Fmain%2Fdocs%2Fcoverage-badge.json)](https://github.com/takara-ai/miru-code/actions/workflows/ci.yml) [![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE) [![bun](https://img.shields.io/badge/runtime-bun%201.1%2B-black)](https://bun.sh)
4 
5 
6**Hybrid code search for AI coding agents.** Find code by meaning, not grep.
7 
8Your AI agent finds the code it needs with up to 50% fewer tokens.
9 
10Miru returns the best **chunks** (path, lines, snippet) for questions like *"where is auth middleware configured?"* — plugged directly into Claude Code, Cursor, Copilot, Codex, and 9+ other agents via MCP.
11 
12---
13 
14## How it fits into your agent's workflow
15 
16![How Miru fits into an agent's workflow](docs/miru-workflow-diagram.svg)
17 
18Miru replaces the grep/glob-style search agents fall back on today. Install once, and every connected agent gets `search` and `find_related` MCP tools automatically.
19 
20 
21## Install via a plugin marketplace
22 
23Prefer this when your IDE has a plugin marketplace — no `bun add -g` step, and updates go through the IDE's own plugin flow. Run `miru setup` first if you have not authenticated yet (see [Set up credentials](#set-up-credentials)).
24 
25**Claude Code:**
26 
27```
28/plugin marketplace add takara-ai/miru-code
29/plugin install miru
30```
31 
32Restart Claude Code (or reload plugins) when prompted. Update: `/plugin marketplace update miru` then reinstall. Remove: `/plugin uninstall miru`.
33 
34**Cursor:**
35 
361. Dashboard → Plugins → Team Marketplaces → **Add Marketplace** → Import from Repo → `takara-ai/miru-code`
372. Customize (sidebar) → find `miru` → **Install** → choose project or user scope
38 
39**Any other IDE, or if marketplace install isn't available:** use the CLI install below.
40 
41### What you get from a plugin install vs the CLI
42 
43Plugin installs don't carry full CLI parity — what each one gives you depends on the IDE's own plugin capabilities, not just Miru's packaging:
44 
45| | Claude Code | Codex | Cursor |
46|---|---|---|---|
47| MCP tools (`search`, `locate`, `expand`, `find_related`) | ✅ | ✅ | ❌ *(plugin ships skills + rules only — no MCP entry yet)* |
48| `miru` / Caveman / STE skills | ✅ | ✅ | ✅ |
49| Dedicated sub-agent (`miru:miru-code`) | ✅ | — | — |
50| Benchmark mode toggle (`/plugin configure`) | ✅ | — | — |
51| Credentials | own plugin-scoped dir | own plugin-scoped dir | n/a |
52 
53`—` means the IDE's plugin schema has no equivalent mechanism to port these to (not a packaging gap we can close): Codex's and Cursor's plugin manifests have no `agents` or `userConfig` fields, so the dedicated sub-agent and the benchmark toggle are Claude-Code-only.
54 
55**Credentials are plugin-scoped, not shared with your CLI install.** Claude Code and Codex plugins each store auth in their own IDE-managed data directory (survives plugin updates, removed on uninstall). This means `miru setup` done via the CLI does **not** carry over to a plugin install, or vice versa — each authenticates independently on first use (interactive device-code login bootstraps automatically). If you use both the CLI and a plugin, you'll sign in twice.
56 
57## Install (CLI)
58 
59```bash
60bun add -g @takara-ai/miru-code
61```
62 
63## Set up credentials
64 
65```bash
66miru setup
67```
68 
69Interactive `miru setup` defaults to device-code login and saves the resulting credentials locally. Manual bearer-token entry is still available with `--key`. If credentials are missing, the interactive MCP/plugin path can bootstrap the same device flow automatically on first use.
70 
71```bash
72miru setup --device # explicit device-code login
73miru setup --key YOUR_TOKEN # store a bearer token directly
74miru setup --clear # remove stored credentials
75```
76 
77Miru stores versioned credentials in `credentials.json` and automatically loads or refreshes them for MCP and CLI use. `TAKARA_API_KEY` still overrides stored credentials when set explicitly.
78 
79## Add to your IDE
80 
81```bash
82miru install
83```
84 
85Interactive TUI — **↑↓** move, **space** toggle, **a** all, **enter** confirm. Pick agents and integrations:
86 
87 
88| Integration | What it does |
89| ----------------------------- | ------------------------------------------------------------------- |
90| MCP server | `search`, `locate`, `expand`, and `find_related` tools in the agent |
91| Instructions | Search policy in `CLAUDE.md` / `AGENTS.md` / `GEMINI.md` |
92| Sub-agent | Dedicated `miru-code` agent file |
93| Cursor rules | Always-on `.cursor/rules/miru-code.mdc` (Cursor only) |
94| Old Miru hooks | Removes stale search hooks from older releases |
95| Caveman *(experimental)* | On-demand chat compression skill (`/caveman`) |
96| STE writing *(experimental)* | On-demand clear technical English for docs (`/ste`) |
97 
98 
99Restart the IDE when done.
100 
101```bash
102miru uninstall # remove miru config
103```
104 
105**Supported:** Cursor · Claude Code · Gemini CLI · Kiro · OpenCode · GitHub Copilot · Codex · VS Code · Visual Studio (Windows) · Windsurf / Devin Desktop
106 
107 
108| IDE | MCP | Instructions / rules | Caveman *(experimental)* | STE *(experimental)* |
109| -------------- | ------------------------------------- | ------------------------------- | ----------------------------------- | --------------------------------------- |
110| Cursor | `~/.cursor/mcp.json` | `~/.cursor/rules/miru-code.mdc` | `~/.agents/skills/caveman/SKILL.md` | `~/.agents/skills/ste/SKILL.md` |
111| Claude Code | `~/.claude.json` | `~/.claude/CLAUDE.md` | `~/.claude/skills/caveman/SKILL.md` | `~/.claude/skills/ste/SKILL.md` |
112| Gemini CLI | `~/.gemini/settings.json` | `~/.gemini/GEMINI.md` | `~/.agents/skills/caveman/SKILL.md` | `~/.agents/skills/ste/SKILL.md` |
113| Kiro | `~/.kiro/settings/mcp.json` | `~/.kiro/steering/miru.md` | `~/.kiro/skills/caveman/SKILL.md` | `~/.kiro/skills/ste/SKILL.md` |
114| OpenCode | `$XDG_CONFIG_HOME/opencode/opencode.json(c)` *(else `~/.config/opencode/…`)* | `…/AGENTS.md` | `~/.agents/skills/caveman/SKILL.md` | `~/.agents/skills/ste/SKILL.md` |
115| GitHub Copilot | `~/.copilot/mcp-config.json` | — | `~/.agents/skills/caveman/SKILL.md` | `~/.agents/skills/ste/SKILL.md` |
116| Codex | `~/.codex/config.toml` | `~/.codex/AGENTS.md` | `~/.agents/skills/caveman/SKILL.md` | `~/.agents/skills/ste/SKILL.md` |
117| VS Code | `…/Code/User/mcp.json` | — | `~/.agents/skills/caveman/SKILL.md` | `~/.agents/skills/ste/SKILL.md` |
118| Visual Studio | `%USERPROFILE%\.mcp.json` | — | `~/.agents/skills/caveman/SKILL.md` | `~/.agents/skills/ste/SKILL.md` |
119| Windsurf | — | — | `~/.agents/skills/caveman/SKILL.md` | `~/.agents/skills/ste/SKILL.md` |
120 
121 
122### Plugin packaging
123 
124- Codex: `.codex-plugin/plugin.json`, `.agents/plugins/marketplace.json`, shares the root `mcp.json`
125- Claude Code: `.claude-plugin/plugin.json`, `.claude-plugin/marketplace.json`, its own `.claude-plugin/mcp.json` (needed for the `userConfig` benchmark toggle — see [What you get from a plugin install vs the CLI](#what-you-get-from-a-plugin-install-vs-the-cli))
126- Cursor: `plugin.json` and `.cursor/rules/miru-code-search.mdc`
127 
128Current limitation:
129 
130- these plugin manifests launch the published Miru runtime through `bunx @takara-ai/miru-code@latest mcp`
131- that means local source edits do not affect plugin behavior until a package version is published
132- and a fully self-contained “no Bun required” plugin install is still future work
133 
134### On-demand skills
135 
136Sub-agent files are also written where supported (see `miru install` plan). Windsurf hooks only *(experimental)* — no MCP entry yet. Caveman is an on-demand Agent Skill (default off): invoke with `/caveman` or “talk like caveman”; stop with “normal mode”. Invocation UI varies by IDE (`/caveman`, `$caveman`, `@caveman`, etc.). Most IDEs share `~/.agents/skills/caveman/SKILL.md` (including Copilot / VS Code / Visual Studio); Claude Code and Kiro keep vendor-native skill dirs. Ownership is tracked on the shared path so uninstalling one IDE keeps the skill while another still owns it; selecting all owners removes it once. STE is an on-demand Agent Skill (default off): invoke with `/ste` or “de-slop this”; keep articles and complete sentences. Most IDEs share `~/.agents/skills/ste/` the same way (ownership via `miru-owners.json`); Claude Code and Kiro keep vendor-native STE dirs.
137 
138<details>
139<summary>STE writing <em>(experimental)</em></summary>
140 
141STE helps write **clear technical English** for docs, runbooks, errors, and release notes (pragmatic ASD-STE100-inspired rules). Invoke with `/ste` or “de-slop this”. Keep articles and complete sentences — not telegraph-style omission.
142 
143**Not ASD-certified.** Full dictionary compliance needs the official standard at [asd-ste100.org](https://www.asd-ste100.org). Miru does not ship the copyrighted ASD dictionary.
144 
145Not for marketing or brand copy. Off by default; enable STE in the installer for any supported IDE. Restart the IDE (or reload skills) after install. For Codex, install also sets `[features] skills = true` in `~/.codex/config.toml` (same as Caveman).
146 
147</details>
148 
149<details>
150<summary>Caveman mode <em>(experimental)</em></summary>
151 
152Caveman compresses **live chat replies** (less filler, max meaning). Intensities: `/caveman lite|full|ultra` (default full). Persisted artifacts (commits, PRs, customer docs) stay normal prose unless you ask otherwise.
153 
154Security / destructive warnings use clear normal prose (auto-clarity) — brevity never hides risk. Session token savings vary; the skill itself costs input tokens. No guaranteed %.
155 
156Off by default at install time. Enable Caveman in the installer for any supported IDE. Restart the IDE (or reload skills) after install. For Codex, the installer also sets `[features] skills = true` in `~/.codex/config.toml` (required for Codex to load skills).
157 
158</details>
159 
160**Team sub-agent in a repo** (optional):
161 
162```bash
163miru init --agent claude --force
164```
165 
166## Try it
167 
168Meaning-based questions → `search`. Exact strings (env vars, symbols, error codes) → `locate`.
169 
170```bash
171miru search "auth middleware" ./src
172miru locate REDIS_HOST ./src
173miru expand src/auth.ts 42 ./src
174miru find-related src/auth.ts 42 ./src
175```
176 
177Terminal output is human-readable; use `--json` for scripts. One-off without installing:
178 
179```bash
180bunx @takara-ai/miru-code@latest search "auth middleware" ./src
181```
182 
183---
184 
185## MCP tools
186 
187When wired via `miru install`, the MCP server exposes `search`, `locate`, `expand`, and `find_related`. `read_benchmark` appears only in [benchmark mode](#benchmark-mode). `repo` is optional: it defaults to the server's startup directory. Set it for another repo, or if your MCP client starts elsewhere. The index is built on the first call and cached for the session.
188 
189 
190| Tool | When to use |
191| ---------------- | --------------------------------------------------------------------------------------- |
192| `search` | Default for code exploration — hybrid semantic + keyword search. One call per question. |
193| `locate` | Exact substrings (env vars, symbols, error codes) — prefer over Grep. |
194| `expand` | More context in the **same file** when a hit has `truncated: true`. |
195| `find_related` | Similar code in **other files** from a `file_path` + `anchor_line`. |
196| `read_benchmark` | Cumulative token-savings rollup *(benchmark mode only)*. |
197 
198 
199### Workflow
200 
2011. **`search`** with `query` — returns compact snippets (~±15 lines) and relevance scores.
2022. If a hit has **`truncated: true`**, call **`expand`** with `file_path` and `anchor_line` — not another search or a full-file read.
2033. To trace similar patterns elsewhere, call **`find_related`** with the same `file_path` and `anchor_line`.
2044. Use your editor's **Read** on `absolute_path` only when editing or when `expand` still lacks context.
205 
206Keep `repo` on follow-up calls when set.
207 
208Prefer these tools over Grep, Glob, or SemanticSearch when Miru MCP is connected — hooks and instructions enforce that when enabled.
209 
210Local repo hits include `absolute_path` for one-click navigation. Parameter reference is below under [MCP parameters](#mcp-parameters).
211 
212## How it works
213 
214Hybrid search: Takara embeddings + BM25 + fusion + reranking. Index **code**, **docs**, **config**, or **all** with `--content`.
215 
216MCP watches local files and updates the index incrementally. Package upgrades invalidate stale caches via the version epoch.
217 
218| OS | Index cache |
219|----|-------------|
220| macOS | `~/Library/Caches/miru` |
221| Linux | `~/.cache/miru` |
222| Windows | `%LOCALAPPDATA%\miru\Cache` |
223 
224### Chunking & languages
225 
226Miru chunks source in tiers: **AST** (tree-sitter, default) → **structural** heuristics → **line** splits.
227 
228**AST chunking** — 26 languages (syntax-aware boundaries via vendored `web-tree-sitter` grammars):
229 
230| Language | Typical extensions |
231|----------|-------------------|
232| astro | `.astro` |
233| bash | `.sh`, `.bash`, `.zsh` |
234| c | `.c` |
235| cpp | `.cpp`, `.h`, `.hpp`, etc. |
236| csharp | `.cs` |
237| css | `.css` |
238| dart | `.dart` |
239| elixir | `.ex`, `.exs` |
240| embeddedtemplate | `.erb`, `.ejs` |
241| go | `.go` |
242| haskell | `.hs` |
243| html | `.html`, `.htm` |
244| java | `.java` |
245| javascript | `.js`, `.jsx`, `.mjs`, `.cjs` |
246| json | `.json` |
247| ocaml | `.ml`, etc. |
248| php | `.php` |
249| python | `.py`, `.pyi` |
250| ruby | `.rb` |
251| rust | `.rs` |
252| scala | `.scala` |
253| solidity | `.sol` |
254| sql | `.sql` |
255| svelte | `.svelte` |
256| typescript | `.ts`, `.tsx`, `.mts`, `.cts` |
257| vue | `.vue` |
258 
259**Structural fallback** (brace/indent heuristics when AST is unavailable): python, go, typescript, javascript, cpp, c.
260 
261**Line fallback:** everything else that gets indexed (kotlin, swift, etc.) — still searchable, coarser chunks.
262 
263Set `MIRU_AST_CHUNKING=0` to disable AST and use structural → lines only.
264 
265## CLI reference
266 
267Run `miru` in a terminal or `miru -h` for the command list, or `miru <command> -h` for details.
268 
269| Command | Purpose |
270| ----------------------------------------- | --------------------------------------------------------------------------------- |
271| `miru setup` | Authenticate and store credentials |
272| `miru install` | Configure IDE (global) |
273| `miru uninstall` | Remove IDE config |
274| `miru search <query> [path]` | Search (`-k N`, `--content`, `--json`) |
275| `miru locate <literal> [path]` | Exact substring in the index |
276| `miru expand <file> <line> [path]` | Adjacent chunks in the same file |
277| `miru find-related <file> <line> [path]` | Related chunks |
278| `miru benchmark on/off/status/clear` | Toggle MCP benchmark mode / clear report |
279| `miru init --agent <id>` | Project-local sub-agent |
280| `miru clear [path]` | Drop index cache (use after big CLI-only refactors) |
281| `miru mcp` | Start MCP server (`--benchmark` for comparisons) |
282 
283 
284CLI uses hyphens (`find-related`); MCP tool names use underscores (`find_related`).
285 
286## Library
287 
288```bash
289bun add @takara-ai/miru-code
290```
291 
292```ts
293import { MiruIndex } from "@takara-ai/miru-code";
294 
295const index = await MiruIndex.fromPath("./src");
296const results = await index.search({ query: "BM25 tokenize" });
297```
298 
299## Environment
300 
301 
302| Variable | Notes |
303| ----------------------------- | ------------------------------------------------------------------------ |
304| `TAKARA_API_KEY` | Required |
305| `MIRU_OPENAI_BASE_URL` | Default `https://infer.takara.ai/v1` |
306| `MIRU_OPENAI_EMBEDDING_MODEL` | Default `ds1-miru-int8` |
307| `MIRU_WORKSPACE_ROOT` | Optional: restrict MCP local `repo` paths to this directory |
308| `MIRU_MAX_INDEX_FILES` | Cap files indexed per operation |
309| `MIRU_ALLOW_HTTP_GIT` | Set `1` to allow plain `http://` git clones |
310| `MIRU_MCP_WATCH` | Set `0` to disable MCP file watch |
311| `MIRU_AST_CHUNKING` | Set `0` to disable tree-sitter AST chunking |
312| `MIRU_BENCHMARK_HISTORY_PATH` | Override; see [Benchmark mode](#benchmark-mode) |
313| `MIRU_CACHE_HOME` | Override the platform cache root (see [How it works](#how-it-works)) |
314| `MIRU_QUIET` | Set `1` to skip the framed CLI banner (subtitle only on color terminals) |
315| `NO_COLOR` | Disable CLI colors |
316 
317 
318See `.env.example` for more.
319 
320## Privacy and API usage
321 
322Miru sends **file contents** to the [Takara inference API](https://takara.ai) when building an index and when embedding search queries. Chunks from your repo are transmitted over HTTPS to generate embeddings. API usage may incur cost depending on your Takara plan.
323 
324If you index proprietary code, make sure that sending snippets to Takara's endpoint fits your security and compliance requirements. `MIRU_WORKSPACE_ROOT` is an opt-in boundary for MCP local `repo` paths only, and restricts indexing to a single workspace directory when set.
325 
326Enterprise self-hosted embeddings (no Takara egress): see [docs/self-hosted-sagemaker.md](docs/self-hosted-sagemaker.md).
327 
328## Benchmark mode
329 
330Optional measurement of how many tokens Miru saves versus a simple Grep workflow (ripgrep + reading the top matched file). Useful when evaluating Miru; leave it off day-to-day. Only local repo paths are compared — git URL repos skip the comparison and return `benchmark_skipped: "local_repo_only"`.
331 
332```bash
333miru benchmark on # add --benchmark to installed MCP configs
334miru benchmark status
335miru benchmark off # prefer off when finished measuring
336miru benchmark clear # delete the global report
337```
338 
339Restart agents after changing mode. `miru install` keeps `--benchmark` if it was already enabled.
340 
341While on, each `search` / `locate` response includes a compact `benchmark` block (`save_pct`, `miru_tok`, `grep_tok`, `saved_tok`, `rank1`). Call `read_benchmark` for a cumulative rollup (or ask the agent when you want totals).
342 
343History is **global** (not per-repo) under Miru's state directory:
344 
345 
346| OS | Default path |
347| ------- | ---------------------------------------------------------------------------- |
348| macOS | `~/Library/Application Support/miru/benchmark-history.json` |
349| Linux | `~/.config/miru/benchmark-history.json` (`$XDG_CONFIG_HOME/miru/…` when set) |
350| Windows | `%APPDATA%\miru\benchmark-history.json` |
351 
352 
353Override with `MIRU_BENCHMARK_HISTORY_PATH`. Append-only JSONL of compact token deltas (no query text); `read_benchmark` returns cumulative totals. Stored in plaintext — `miru benchmark clear` or uninstall on shared machines.
354 
355## MCP parameters
356 
357`**search**`
358 
359 
360| Param | Required | Notes |
361| ---------------- | -------- | --------------------------------------- |
362| `query` | yes | Natural language or code query |
363| `repo` | no | Startup directory by default; set for another repo |
364| `include` | no | Gitignore-style glob patterns; only matching files are searched (same as `locate.include`) |
365| `exclude` | no | Gitignore-style glob patterns; matching files are skipped (same as `locate.exclude`) |
366| `dedupe_by_file` | no | Keep best hit per file (default `true`) |
367 
368 
369`**locate**`
370 
371 
372| Param | Required | Notes |
373| ------------- | -------- | ----------------------------------------------------------------------------------- |
374| `literal` | yes | Exact substring to find |
375| `repo` | no | Startup directory by default; set for another repo |
376| `include` | no | Gitignore-style glob patterns; only matching files are searched |
377| `exclude` | no | Gitignore-style glob patterns; matching files are skipped |
378| `mode` | no | `count` · `locations` · `lines` (default). Prefer `count`/`locations` when possible |
379| `limit` | no | Cap returned hits. Omit to return all matches |
380| `ignore_case` | no | Case-insensitive match (default `false`) |
381 
382 
383`**expand**`
384 
385 
386| Param | Required | Notes |
387| ------------------ | -------- | --------------------------------------------------------------------- |
388| `file_path` | yes | From hit `file_path` or `absolute_path` (local repos) |
389| `anchor_line` | yes | From the search hit (`anchor_line` when truncated, else `start_line`) |
390| `repo` | no | Same repo as the search, if set |
391| `before` / `after` | no | Extra chunks before/after anchor (default 1 each) |
392 
393 
394`**find_related**`
395 
396 
397| Param | Required | Notes |
398| ------------- | -------- | -------------------------------------------- |
399| `file_path` | yes | From a search hit |
400| `anchor_line` | yes | From the search hit |
401| `repo` | no | Same repo as the search, if set |
402 
403 
404`**read_benchmark**` *(benchmark mode only)*
405 
406 
407| Param | Required | Notes |
408| ------ | -------- | ---------------------------------------------------------------------- |
409| `repo` | no | Filter rollup to one local path or git URL. Omit for all saved queries |
410 
411 
412## Manual MCP (skip `miru install`)
413 
414```json
415{
416 "miru": {
417 "command": "miru",
418 "args": ["mcp"]
419 }
420}
421```
422 
423For benchmark mode, set `"args": ["mcp", "--benchmark"]` (or append the flag). Prefer `miru benchmark on` after a normal install — it updates every agent config.
424 
425Run `miru setup` once so the server can load credentials from `credentials.json`. If the MCP server starts in an interactive terminal without stored credentials, it will start device login automatically.
426 
427Use `bunx` + `@takara-ai/miru-code@latest` if `miru` is not global. The installer uses this command so each new MCP server launch can pick up a published version without a global update. A running server keeps its current version until restarted. Wrapper key varies by IDE (`mcpServers`, `servers`, or `mcp`).
428 
429Older headless MCP configs that launch `miru` without a subcommand, with or without MCP flags, continue to work. Run `miru install` again to update managed entries to the explicit `mcp` subcommand, then restart your coding agent. For manual MCP configs, add `mcp` after the executable or package name.
430 
431## Developing
432 
433```bash
434git clone https://github.com/takara-ai/miru-code.git && cd miru-code
435bun install && cp .env.example .env.local
436bun test && bun run typecheck
437```
438 
439See [CONTRIBUTING.md](CONTRIBUTING.md) for pre-commit hooks, commit message conventions, and the PR process.
440 
441Local MCP: `"command": "bun", "args": ["/path/to/miru-code/src/cli.ts"]`
442 
443`miru -h` / setup print a framed wordmark on color terminals (`MIRU_QUIET=1` for subtitle only). Crane art lives in `src/brand-banner.ts`; regenerate with `bun run scripts/render-crane-art.ts` (ImageMagick). The crane is a registered mark of Takara.ai Ltd.
444 
445## Credits
446 
447Miru uses work by [MinishLab](https://github.com/MinishLab). Thank you to its authors.
448 
449- **[semble](https://github.com/MinishLab/semble)** (MIT). Miru ports parts of semble. These parts include the file walker, the index layout, the BM25 tokenizer, the hybrid search pipeline, and the ranking signals.
450- **[potion-code](https://huggingface.co/minishlab/potion-code-16M)** and **[Model2Vec](https://github.com/MinishLab/model2vec)** (MIT). The WordPiece tokenizer in `tokenizer/tokenizer.json` comes from potion-code.
451 
452See [NOTICE](./NOTICE) for the licence text and the Model2Vec citation.
453 
454## License
455 
456MIT
457 

Discussion

Alternatives