Code migration skill
Move code from one implementation to another function by function, tests first and code second, and check two implementations of one app for parity, with jscpd --compare as the progress measure and a coverage map binding each function to its tests.
by kucherenko·MIT license·★ 6,331 Stars on the repo·GitHub ↗
npx degit kucherenko/jscpd/skills/code-migration#master ~/.claude/skills/code-migrationChecked ·commit master
Files of Code migration
Show the full text201 lines
code-migration
jscpd --compare SOURCE TARGET pairs every function of one folder with the function of the other folder that does the same job, in any pair of languages, and lists the functions that have no counterpart. This skill uses it to measure a port: what is ported, what is left, and whether the port you just wrote was recognized.
A port runs in two phases: the tests first, then the code they check. A coverage report of the source's tests tells which tests exercise which function, so each function is ported together with the tests that prove it works. See Plan: tests first, then code.
Two words are used throughout:
- source: the implementation you port from, such as the iOS app, the Python library, or the JavaScript package;
- target: the implementation you port to, such as the Android app, the Rust crate, or the TypeScript rewrite.
Neither has to be old or new. The source may stay in production and keep changing, and the target may already hold features the source lacks. For two implementations that both live on (ios/ and android/), the same report shows parity; see Parity.
jscpd finds functions in JavaScript, TypeScript, JSX, TSX, Vue, Svelte, Astro, Python, Rust, Go, Java, Kotlin, C#, C, C++, PHP, Ruby, Scala and Swift. The compare-codebases skill explains how the comparison pairs functions and how to check its result; the jscpd skill covers the rest of the tool.
Setup
--compare pairs functions with a code embedding model that runs inside jscpd. Download it once (548 MB, CodeRankEmbed). Ask the user before you start the download:
npx jscpd --semantic-download
jscpd caches the vectors, so a repeat run embeds only the functions whose code changed and takes seconds.
Pick the two paths and keep them fixed for the whole port: the source first, the target second. Two paths are required, and they must not overlap (app/ and app/android/ is refused). The target may be empty at the start.
One run measures both phases: the report has a Code block and a Tests block (a code and a tests section in JSON), and a test pairs only with a test. jscpd tells a test by the conventions of its language: test files such as *_test.go, test_*.py, *.test.ts or *Test.java, folders such as tests/, __tests__/ or src/test/, Rust tests in #[cfg(test)] modules, and JavaScript test cases such as it('rounds cents', () => …). Phase 1 reads the Tests block, phase 2 the Code block.
Keep vendored and generated code out
--compare counts and embeds every function in both paths. Vendored dependencies (vendor/ from cargo vendor or go mod vendor, third_party/), installed packages (node_modules/, .venv/), build output (target/, build/, dist/) and generated code are not part of the port, yet a vendored crate tree alone holds thousands of functions. With them in a path, the first run embeds all of them and takes tens of minutes instead of seconds, and the percentages describe the dependencies instead of the port.
jscpd skips what .gitignore excludes, but only inside a git repository. Before the first run:
Check both paths for such folders. A target you create for the port gets them as soon as you build it or vendor its dependencies, so check again after the first build.
Suggest a
.gitignoreto the user that lists the folders the target's language and tools produce, plus the report folder. For a Rust addon built with napi-rs:/target/ /vendor/ node_modules/ *.node .jscpd-compare/If the target is not in a git repository, tell the user that jscpd reads the
.gitignoreonly aftergit init.Until the
.gitignoreworks, pass the same folders with--ignore, and keep the globs the same for every run:npx jscpd --compare node-lib/ rust-lib/ --ignore "**/vendor/**,**/target/**,**/node_modules/**"
The totals line shows when something slipped through. Suspect vendored or generated code in a path when the target has far more functions than the source, or when a run embeds hundreds of functions after a small change.
Measure
Run the console report for yourself and for the user. Here a Python billing library is being ported to TypeScript:
npx jscpd --compare billing-py/ billing-ts/
71% 5 of 7 functions in billing-py/ have a counterpart in billing-ts/
80% 4 of 5 functions in billing-ts/ have a counterpart in billing-py/
billing-py/
file paired similarity counterpart
billing.py 4 / 5 0.89 billing.ts
shipping.py 1 / 2 0.91 shipping.ts
billing-ts/
file paired similarity counterpart
billing.ts 3 / 4 0.89 billing.py
shipping.ts 1 / 1 0.91 shipping.py
Paired under other names (1):
billing-py/ billing-ts/ similarity
billing.py:28 tax_for_region billing.ts:27 salesTax 0.87 high
Only in billing-py/ (2):
billing.py (1)
46 due_date 6 lines
shipping.py (1)
18 estimate_delivery_days 8 lines
Only in billing-ts/ (1):
billing.ts (1)
39 toCurrency 8 lines
- The first line is the port's progress: the share of the source's functions that have a counterpart in the target. "Only in" the source is the work left.
- The second line and "Only in" the target describe the target's own code: helpers the port needed, features the source never had, or a port jscpd did not recognize (see Misses).
- Each file row gives its paired functions, the mean similarity of their pairs (with the number of
lowpairs, if any), and the counterpart file: the file on the other side that holds most of its counterparts. Use the counterpart to decide where a missing function goes. - "Paired under other names" lists the pairs whose names differ even once case and underscores are ignored: renamed ports, constructors, platform names. Each has its similarity and level (see Reading pairs). These are ported already, even though a search by name would not find them.
- With an empty target the report is one line of totals plus
billing-ts/ has no functions yet.
For work you plan and track, read the JSON report instead of the console:
npx jscpd --compare billing-py/ billing-ts/ -r json -o .jscpd-compare --silent
jscpd measures tests and code apart and pairs a test only with a test, so .jscpd-compare/jscpd-compare.json has a code and a tests section of the same shape. Each section has sides[0] (the source) and sides[1] (the target), each with path, functions, matched, percentage, files (file, functions, matched, counterpart, similarity, lowPairs) unmatched (file, name, start, end) and readyToPort (the unmatched functions whose callees all have a counterpart, with their number of callers, most called first), and pairs, each with a (source), b (target), similarity, level (high, medium, low), renamed (the names differ) and matchedBy (code or name). Paths are relative to each side's folder. Add .jscpd-compare/ to .gitignore or write the report outside the repository.
-r console-full also prints every pair with its similarity, -r markdown writes jscpd-compare.md, a table you can paste into a PR description or a tracking issue, and -r html writes jscpd-compare.html, a migration map the user can open to see the two dependency graphs side by side, or the same pairs as a table.
Plan: tests first, then code
Port the tests before the code. Ported tests define, before any target code exists, what it means for each function to be ported: when the function lands, its tests pass or show what still differs. A port written before its tests can only be checked against your reading of the source.
Bind tests to functions with coverage
A test's name does not say which functions it runs: a test of checkout also runs apply_discount and tax_for_region. Coverage does. Before porting anything, build a map from each source function to the tests that exercise it.
- Get a coverage report of the source's tests that says which test ran which lines: recorded per test, or per test file when the project's tooling has no per-test mode.
- Take the source's functions and their lines from the
codesection of the JSON report. Every counted function is either in the source'sunmatchedlist or theaside of a pair, each withfile,name,startandend. Functions under--min-tokensor--min-linesare not listed there; take those from the coverage report's own function list, which most formats have. - A test covers a function when it runs a line from the function's
startto itsend. Write the map both ways, function to tests and test to functions, to a file next to the report (for example.jscpd-compare/test-map.json), and rebuild it when the source's tests change.
Use the map for three things:
- which tests to port with a function, and which functions a ported test needs before it can pass;
- the source functions no test covers. Before porting such a function, write a test for it on the source side that records what it does today (a characterization test), and port that test with the others. A port with no test has nothing to check it;
- the order of work: tests that need few functions come first, since they turn green soonest.
Coverage says which code a test runs, not what it checks. A function that a test only passes through on the way to another is weakly bound to that test; prefer the tests whose names or assertions are about the function.
Phase 1: port the tests
- Take the source's unmatched tests from the
testssection of the JSON report. A JavaScript or TypeScript test case written as a callback,it('rounds cents', () => …), goes by its title, so it pairs withtest_rounds_centsin pytest orrounds_centsin Rust like any named test. - Port the tests file by file, into the target's test layout, against the API the target will have. Keep inputs and expected values exactly as they are. Where the target must behave differently (a platform limit, a language's number types), write the difference into the test with a comment and tell the user.
- A ported test calls functions that do not exist yet, so it fails. In a compiled language it stops the whole test build. Keep such tests out of the build until their functions land, for example a module declaration left commented out, a cfg feature or an excluded source set, and turn them on one by one in phase 2. Do not write empty functions to make the tests compile: a stub under a source name can pair by name and count as ported.
- Run the target's tests. The ported tests that pass already cover what the target has; the failing or disabled ones, read against the test map, are the work list for phase 2.
Phase 2: port the code
Port the code one function at a time, with its tests already in place.
- Run the JSON report and take the source's
unmatchedlist from itscodesection. - Order the work. Start with functions whose tests are ported, and port a function after the functions it calls, since a caller ported before its helpers has nothing to call. The source's
readyToPortlist is that order's front: the functions whose callees all have a counterpart already, most called first. Take the next one from there, and rerun the report after each batch, since every port can make its callers ready. Within that order, finish one file before starting the next, so each target file fills up in one go. - Before writing a port, make sure the target does not have it already. Pairs with
renamed: trueare ported functions under other names; they are not inunmatched, but check them when the user asks where a function went. Then search the target for an equivalent that jscpd missed: in the counterpart file fromfiles, under other names (constructors, merged functions, a platform or library call that replaces the helper). If one exists, do not port the function again. Note the equivalent and move on. - Write the port in the counterpart file, or in the file the target's layout puts it in. Keep the name recognizable in the target language's convention (
encodeBinarybecomesencode_binaryin Rust or Python), because jscpd pairs short functions by name. Follow the style of pairs that are already done in the same file.console-fulllists them. - Turn on the function's ported tests, as the test map lists them, and run them. When they fail, fix the port, not the test, unless the user agrees that the target should behave differently. A pair in the report says the two functions look alike; the tests say whether they behave alike.
- Run the target's coverage for those tests and check that they run the new function. A ported test that never reaches it is bound to something else on the target (a helper, a mock, an old path), and it proves nothing about the port.
- Run the comparison again with the same paths and options. Check that the function has left
unmatched, that its pair joins the right function in the target, and its level. Ahighpair is done. Alowpair to your port usually means the port differs a lot from the source in structure; read both and make sure the behavior matches. A pair to a different function means the port is not recognized, or it resembles the wrong thing; read both before going on. - Report progress to the user with both measures, as the reports print them: tests ported and passing, and functions ported (
functions 73% → 76%, 3 ported: …; tests 41 of 44 ported, 38 passing), with the functions you decided not to port and why.
Repeat until the source's unmatched list holds only functions you decided not to port. Typical reasons: dead code (check with npx jscpd --dead-code on the source where it supports the language), code the target platform or its libraries provide (a JSON parser, a retry helper), platform glue with no equivalent on the target (an iOS delegate callback, an Android notification channel), and features the user decided not to carry over. List these in a notes file or the PR, one line each, so the remaining percentage is explained.
When the source keeps changing during the port, a function that was paired can come back as unmatched after a rewrite, and new source functions join the list. Rerun the report before each batch of work instead of reusing an old list.
Parity between two implementations
When both implementations live on, as the iOS and the Android app do, npx jscpd --compare ios/ android/ shows parity. Read both "Only in" lists:
- a function on one side only is either platform code (Android notification channels, iOS delegate callbacks) or a feature the other side lacks. Report each missing feature to the user before implementing it;
- to add a missing feature, treat the side that has it as the source for that feature: bind its tests with coverage, port the tests, then the code, as in Plan: tests first, then code.
jscpd pairs functions within modules that already match. A module is the folder right under the deepest folder each side's files share, so keep the two trees parallel (ios/<plugin> and android/<plugin>) for the best pairing.
Reading pairs
levelrates the similarity on the scale of the model in use, so it means the same with every model.high(0.7125 and up with the default CodeRankEmbed, across languages) is almost always the same function.mediumis usually the same function, restructured.lowneeds reading both functions, since related code pairs there too (a function counting UTF-8 bytes paired with one converting a string to them). On the Tauri plugins, everylowpair joined two differently named functions, and one of them joined two different plugins.matchedBy: codemeans the two functions are each other's closest match and the similarity stands out.matchedBy: namemeans the names match once case and underscores are ignored, and the code is similar enough. It catches short ports. Read these pairs, because a same-named function can do something else.- The similarity does not see small differences in behavior. Two versions that drifted apart still pair. Only tests and reading the code catch drift.
Misses
jscpd compares functions only: types, constants, enums with data, SQL and UI markup (storyboards, Android layouts) are not in the report, so port and check them yourself. Known gaps in pairing:
- constructors across languages (a Java constructor and Rust's
new, Kotlin'sconstructor, Swift'sinit) pair only when their code is similar enough; - a short function renamed in the port (
add_historyfor_finder_penalty_add_history); - one function split into several, or several merged into one: the report may pair only the closest part and list the rest as unmatched;
- anonymous functions (callbacks, closures) take no part at all, except JavaScript and TypeScript test cases such as
it('rounds cents', () => …), which go by their titles.
A function listed only in the target that you know is a port of a source function is one of these. Do not rename working code only to raise the number, and never add stubs or empty functions with source names: jscpd may pair a stub by name, and the progress would then report work that was not done.
Options
--min-tokens(30 with--compare) and--min-lines(5) decide which functions count toward the totals. Smaller functions still pair as partners. Raise them to focus on substantial functions, and keep them fixed across runs so percentages compare.--ignoreleaves files out, such as vendored dependencies, build output, generated code or fixtures (see Keep vendored and generated code out);--patternnarrows the comparison to some files, such as the tests of one module. Keep the globs fixed across runs.--formatnarrows the walk to some languages, for example when the source mixes the code being ported with build scripts.--semantic-modeland--semantic-urlpick another embedding model or an OpenAI-compatible API (npx jscpd --semantic-modelslists the models with calibrated thresholds). Keep the model fixed across runs, because each model scores on its own scale.- The exit code is 0 whatever the progress, so the command does not fail a build on its own.
Rules
- Keep the two paths, their order, the options and the model the same from run to run, or the numbers stop being comparable.
- Keep vendored dependencies, installed packages and build output out of both paths: through a
.gitignorein a git repository, which you suggest to the user, or with--ignore. - Port the tests before the code, and a function only together with the tests the coverage map binds to it.
- Treat the report as a map of what to read and what is left. It does not prove the port is correct; tests do.
- Never game the percentage: no stubs, no renames made only for jscpd, no deleting source functions to shrink the denominator without the user's decision.
- When the report and your reading of the code disagree, trust the code and tell the user which pair looked wrong.
| 1 | |
| 2 | name code-migration |
| 3 | description Move code from one implementation to another function by function, tests first and code second, and check two implementations of one app for parity, with jscpd --compare as the progress measure and a coverage map binding each function to its tests. Use when porting a library or app to another language or framework (Java to Kotlin, JavaScript to Rust, a Python library to TypeScript, an iOS app to Android), when asked what is left to port, or when comparing the Android and iOS versions of an app. |
| 4 | |
| 5 | |
| 6 | # code-migration |
| 7 | |
| 8 | `jscpd --compare SOURCE TARGET` pairs every function of one folder with the function of the other folder that does the same job, in any pair of languages, and lists the functions that have no counterpart. This skill uses it to measure a port: what is ported, what is left, and whether the port you just wrote was recognized. |
| 9 | |
| 10 | A port runs in two phases: the tests first, then the code they check. A coverage report of the source's tests tells which tests exercise which function, so each function is ported together with the tests that prove it works. See [Plan: tests first, then code]. |
| 11 | |
| 12 | Two words are used throughout: |
| 13 | |
| 14 | **source**: the implementation you port from, such as the iOS app, the Python library, or the JavaScript package; |
| 15 | **target**: the implementation you port to, such as the Android app, the Rust crate, or the TypeScript rewrite. |
| 16 | |
| 17 | Neither has to be old or new. The source may stay in production and keep changing, and the target may already hold features the source lacks. For two implementations that both live on (`ios/` and `android/`), the same report shows parity; see [Parity]. |
| 18 | |
| 19 | jscpd finds functions in JavaScript, TypeScript, JSX, TSX, Vue, Svelte, Astro, Python, Rust, Go, Java, Kotlin, C#, C, C++, PHP, Ruby, Scala and Swift. The [compare-codebases] skill explains how the comparison pairs functions and how to check its result; the [jscpd] skill covers the rest of the tool. |
| 20 | |
| 21 | ## Setup |
| 22 | |
| 23 | `--compare` pairs functions with a code embedding model that runs inside jscpd. Download it once (548 MB, CodeRankEmbed). Ask the user before you start the download: |
| 24 | |
| 25 | |
| 26 | npx jscpd --semantic-download |
| 27 | |
| 28 | |
| 29 | jscpd caches the vectors, so a repeat run embeds only the functions whose code changed and takes seconds. |
| 30 | |
| 31 | Pick the two paths and keep them fixed for the whole port: the source first, the target second. Two paths are required, and they must not overlap (`app/` and `app/android/` is refused). The target may be empty at the start. |
| 32 | |
| 33 | One run measures both phases: the report has a `Code` block and a `Tests` block (a `code` and a `tests` section in JSON), and a test pairs only with a test. jscpd tells a test by the conventions of its language: test files such as `*_test.go`, `test_*.py`, `*.test.ts` or `*Test.java`, folders such as `tests/`, `__tests__/` or `src/test/`, Rust tests in `#[cfg(test)]` modules, and JavaScript test cases such as `it('rounds cents', () => …)`. Phase 1 reads the `Tests` block, phase 2 the `Code` block. |
| 34 | |
| 35 | ### Keep vendored and generated code out |
| 36 | |
| 37 | `--compare` counts and embeds every function in both paths. Vendored dependencies (`vendor/` from `cargo vendor` or `go mod vendor`, `third_party/`), installed packages (`node_modules/`, `.venv/`), build output (`target/`, `build/`, `dist/`) and generated code are not part of the port, yet a vendored crate tree alone holds thousands of functions. With them in a path, the first run embeds all of them and takes tens of minutes instead of seconds, and the percentages describe the dependencies instead of the port. |
| 38 | |
| 39 | jscpd skips what `.gitignore` excludes, but only inside a git repository. Before the first run: |
| 40 | |
| 41 | Check both paths for such folders. A target you create for the port gets them as soon as you build it or vendor its dependencies, so check again after the first build. |
| 42 | Suggest a `.gitignore` to the user that lists the folders the target's language and tools produce, plus the report folder. For a Rust addon built with napi-rs: |
| 43 | |
| 44 | |
| 45 | /target/ |
| 46 | /vendor/ |
| 47 | node_modules/ |
| 48 | *.node |
| 49 | .jscpd-compare/ |
| 50 | |
| 51 | |
| 52 | If the target is not in a git repository, tell the user that jscpd reads the `.gitignore` only after `git init`. |
| 53 | Until the `.gitignore` works, pass the same folders with `--ignore`, and keep the globs the same for every run: |
| 54 | |
| 55 | |
| 56 | npx jscpd --compare node-lib/ rust-lib/ --ignore "**/vendor/**,**/target/**,**/node_modules/**" |
| 57 | |
| 58 | |
| 59 | The totals line shows when something slipped through. Suspect vendored or generated code in a path when the target has far more functions than the source, or when a run embeds hundreds of functions after a small change. |
| 60 | |
| 61 | ## Measure |
| 62 | |
| 63 | Run the console report for yourself and for the user. Here a Python billing library is being ported to TypeScript: |
| 64 | |
| 65 | |
| 66 | npx jscpd --compare billing-py/ billing-ts/ |
| 67 | |
| 68 | |
| 69 | |
| 70 | 71% 5 of 7 functions in billing-py/ have a counterpart in billing-ts/ |
| 71 | 80% 4 of 5 functions in billing-ts/ have a counterpart in billing-py/ |
| 72 | |
| 73 | billing-py/ |
| 74 | file paired similarity counterpart |
| 75 | billing.py 4 / 5 0.89 billing.ts |
| 76 | shipping.py 1 / 2 0.91 shipping.ts |
| 77 | |
| 78 | billing-ts/ |
| 79 | file paired similarity counterpart |
| 80 | billing.ts 3 / 4 0.89 billing.py |
| 81 | shipping.ts 1 / 1 0.91 shipping.py |
| 82 | |
| 83 | Paired under other names (1): |
| 84 | billing-py/ billing-ts/ similarity |
| 85 | billing.py:28 tax_for_region billing.ts:27 salesTax 0.87 high |
| 86 | |
| 87 | Only in billing-py/ (2): |
| 88 | billing.py (1) |
| 89 | 46 due_date 6 lines |
| 90 | shipping.py (1) |
| 91 | 18 estimate_delivery_days 8 lines |
| 92 | |
| 93 | Only in billing-ts/ (1): |
| 94 | billing.ts (1) |
| 95 | 39 toCurrency 8 lines |
| 96 | |
| 97 | |
| 98 | The first line is the port's progress: the share of the source's functions that have a counterpart in the target. "Only in" the source is the work left. |
| 99 | The second line and "Only in" the target describe the target's own code: helpers the port needed, features the source never had, or a port jscpd did not recognize (see [Misses]). |
| 100 | Each file row gives its paired functions, the mean similarity of their pairs (with the number of `low` pairs, if any), and the counterpart file: the file on the other side that holds most of its counterparts. Use the counterpart to decide where a missing function goes. |
| 101 | "Paired under other names" lists the pairs whose names differ even once case and underscores are ignored: renamed ports, constructors, platform names. Each has its similarity and level (see [Reading pairs]). These are ported already, even though a search by name would not find them. |
| 102 | With an empty target the report is one line of totals plus `billing-ts/ has no functions yet`. |
| 103 | |
| 104 | For work you plan and track, read the JSON report instead of the console: |
| 105 | |
| 106 | |
| 107 | npx jscpd --compare billing-py/ billing-ts/ -r json -o .jscpd-compare --silent |
| 108 | |
| 109 | |
| 110 | jscpd measures tests and code apart and pairs a test only with a test, so `.jscpd-compare/jscpd-compare.json` has a `code` and a `tests` section of the same shape. Each section has `sides[0]` (the source) and `sides[1]` (the target), each with `path`, `functions`, `matched`, `percentage`, `files` (`file`, `functions`, `matched`, `counterpart`, `similarity`, `lowPairs`) `unmatched` (`file`, `name`, `start`, `end`) and `readyToPort` (the unmatched functions whose callees all have a counterpart, with their number of `callers`, most called first), and `pairs`, each with `a` (source), `b` (target), `similarity`, `level` (`high`, `medium`, `low`), `renamed` (the names differ) and `matchedBy` (`code` or `name`). Paths are relative to each side's folder. Add `.jscpd-compare/` to `.gitignore` or write the report outside the repository. |
| 111 | |
| 112 | `-r console-full` also prints every pair with its similarity, `-r markdown` writes `jscpd-compare.md`, a table you can paste into a PR description or a tracking issue, and `-r html` writes `jscpd-compare.html`, a migration map the user can open to see the two dependency graphs side by side, or the same pairs as a table. |
| 113 | |
| 114 | ## Plan: tests first, then code |
| 115 | |
| 116 | Port the tests before the code. Ported tests define, before any target code exists, what it means for each function to be ported: when the function lands, its tests pass or show what still differs. A port written before its tests can only be checked against your reading of the source. |
| 117 | |
| 118 | ### Bind tests to functions with coverage |
| 119 | |
| 120 | A test's name does not say which functions it runs: a test of `checkout` also runs `apply_discount` and `tax_for_region`. Coverage does. Before porting anything, build a map from each source function to the tests that exercise it. |
| 121 | |
| 122 | Get a coverage report of the source's tests that says which test ran which lines: recorded per test, or per test file when the project's tooling has no per-test mode. |
| 123 | Take the source's functions and their lines from the `code` section of the JSON report. Every counted function is either in the source's `unmatched` list or the `a` side of a pair, each with `file`, `name`, `start` and `end`. Functions under `--min-tokens` or `--min-lines` are not listed there; take those from the coverage report's own function list, which most formats have. |
| 124 | A test covers a function when it runs a line from the function's `start` to its `end`. Write the map both ways, function to tests and test to functions, to a file next to the report (for example `.jscpd-compare/test-map.json`), and rebuild it when the source's tests change. |
| 125 | |
| 126 | Use the map for three things: |
| 127 | |
| 128 | which tests to port with a function, and which functions a ported test needs before it can pass; |
| 129 | the source functions no test covers. Before porting such a function, write a test for it on the source side that records what it does today (a characterization test), and port that test with the others. A port with no test has nothing to check it; |
| 130 | the order of work: tests that need few functions come first, since they turn green soonest. |
| 131 | |
| 132 | Coverage says which code a test runs, not what it checks. A function that a test only passes through on the way to another is weakly bound to that test; prefer the tests whose names or assertions are about the function. |
| 133 | |
| 134 | ### Phase 1: port the tests |
| 135 | |
| 136 | Take the source's unmatched tests from the `tests` section of the JSON report. A JavaScript or TypeScript test case written as a callback, `it('rounds cents', () => …)`, goes by its title, so it pairs with `test_rounds_cents` in pytest or `rounds_cents` in Rust like any named test. |
| 137 | Port the tests file by file, into the target's test layout, against the API the target will have. Keep inputs and expected values exactly as they are. Where the target must behave differently (a platform limit, a language's number types), write the difference into the test with a comment and tell the user. |
| 138 | A ported test calls functions that do not exist yet, so it fails. In a compiled language it stops the whole test build. Keep such tests out of the build until their functions land, for example a module declaration left commented out, a cfg feature or an excluded source set, and turn them on one by one in phase 2. Do not write empty functions to make the tests compile: a stub under a source name can pair by name and count as ported. |
| 139 | Run the target's tests. The ported tests that pass already cover what the target has; the failing or disabled ones, read against the test map, are the work list for phase 2. |
| 140 | |
| 141 | ### Phase 2: port the code |
| 142 | |
| 143 | Port the code one function at a time, with its tests already in place. |
| 144 | |
| 145 | Run the JSON report and take the source's `unmatched` list from its `code` section. |
| 146 | Order the work. Start with functions whose tests are ported, and port a function after the functions it calls, since a caller ported before its helpers has nothing to call. The source's `readyToPort` list is that order's front: the functions whose callees all have a counterpart already, most called first. Take the next one from there, and rerun the report after each batch, since every port can make its callers ready. Within that order, finish one file before starting the next, so each target file fills up in one go. |
| 147 | Before writing a port, make sure the target does not have it already. Pairs with `renamed: true` are ported functions under other names; they are not in `unmatched`, but check them when the user asks where a function went. Then search the target for an equivalent that jscpd missed: in the counterpart file from `files`, under other names (constructors, merged functions, a platform or library call that replaces the helper). If one exists, do not port the function again. Note the equivalent and move on. |
| 148 | Write the port in the counterpart file, or in the file the target's layout puts it in. Keep the name recognizable in the target language's convention (`encodeBinary` becomes `encode_binary` in Rust or Python), because jscpd pairs short functions by name. Follow the style of pairs that are already done in the same file. `console-full` lists them. |
| 149 | Turn on the function's ported tests, as the test map lists them, and run them. When they fail, fix the port, not the test, unless the user agrees that the target should behave differently. A pair in the report says the two functions look alike; the tests say whether they behave alike. |
| 150 | Run the target's coverage for those tests and check that they run the new function. A ported test that never reaches it is bound to something else on the target (a helper, a mock, an old path), and it proves nothing about the port. |
| 151 | Run the comparison again with the same paths and options. Check that the function has left `unmatched`, that its pair joins the right function in the target, and its level. A `high` pair is done. A `low` pair to your port usually means the port differs a lot from the source in structure; read both and make sure the behavior matches. A pair to a different function means the port is not recognized, or it resembles the wrong thing; read both before going on. |
| 152 | Report progress to the user with both measures, as the reports print them: tests ported and passing, and functions ported (`functions 73% → 76%, 3 ported: …; tests 41 of 44 ported, 38 passing`), with the functions you decided not to port and why. |
| 153 | |
| 154 | Repeat until the source's `unmatched` list holds only functions you decided not to port. Typical reasons: dead code (check with `npx jscpd --dead-code` on the source where it supports the language), code the target platform or its libraries provide (a JSON parser, a retry helper), platform glue with no equivalent on the target (an iOS delegate callback, an Android notification channel), and features the user decided not to carry over. List these in a notes file or the PR, one line each, so the remaining percentage is explained. |
| 155 | |
| 156 | When the source keeps changing during the port, a function that was paired can come back as unmatched after a rewrite, and new source functions join the list. Rerun the report before each batch of work instead of reusing an old list. |
| 157 | |
| 158 | ## Parity between two implementations |
| 159 | |
| 160 | When both implementations live on, as the iOS and the Android app do, `npx jscpd --compare ios/ android/` shows parity. Read both "Only in" lists: |
| 161 | |
| 162 | a function on one side only is either platform code (Android notification channels, iOS delegate callbacks) or a feature the other side lacks. Report each missing feature to the user before implementing it; |
| 163 | to add a missing feature, treat the side that has it as the source for that feature: bind its tests with coverage, port the tests, then the code, as in [Plan: tests first, then code]. |
| 164 | |
| 165 | jscpd pairs functions within modules that already match. A module is the folder right under the deepest folder each side's files share, so keep the two trees parallel (`ios/<plugin>` and `android/<plugin>`) for the best pairing. |
| 166 | |
| 167 | ## Reading pairs |
| 168 | |
| 169 | `level` rates the similarity on the scale of the model in use, so it means the same with every model. `high` (0.7125 and up with the default CodeRankEmbed, across languages) is almost always the same function. `medium` is usually the same function, restructured. `low` needs reading both functions, since related code pairs there too (a function counting UTF-8 bytes paired with one converting a string to them). On the Tauri plugins, every `low` pair joined two differently named functions, and one of them joined two different plugins. |
| 170 | `matchedBy: code` means the two functions are each other's closest match and the similarity stands out. |
| 171 | `matchedBy: name` means the names match once case and underscores are ignored, and the code is similar enough. It catches short ports. Read these pairs, because a same-named function can do something else. |
| 172 | The similarity does not see small differences in behavior. Two versions that drifted apart still pair. Only tests and reading the code catch drift. |
| 173 | |
| 174 | ## Misses |
| 175 | |
| 176 | jscpd compares functions only: types, constants, enums with data, SQL and UI markup (storyboards, Android layouts) are not in the report, so port and check them yourself. Known gaps in pairing: |
| 177 | |
| 178 | constructors across languages (a Java constructor and Rust's `new`, Kotlin's `constructor`, Swift's `init`) pair only when their code is similar enough; |
| 179 | a short function renamed in the port (`add_history` for `_finder_penalty_add_history`); |
| 180 | one function split into several, or several merged into one: the report may pair only the closest part and list the rest as unmatched; |
| 181 | anonymous functions (callbacks, closures) take no part at all, except JavaScript and TypeScript test cases such as `it('rounds cents', () => …)`, which go by their titles. |
| 182 | |
| 183 | A function listed only in the target that you know is a port of a source function is one of these. Do not rename working code only to raise the number, and never add stubs or empty functions with source names: jscpd may pair a stub by name, and the progress would then report work that was not done. |
| 184 | |
| 185 | ## Options |
| 186 | |
| 187 | `--min-tokens` (30 with `--compare`) and `--min-lines` (5) decide which functions count toward the totals. Smaller functions still pair as partners. Raise them to focus on substantial functions, and keep them fixed across runs so percentages compare. |
| 188 | `--ignore` leaves files out, such as vendored dependencies, build output, generated code or fixtures (see [Keep vendored and generated code out]); `--pattern` narrows the comparison to some files, such as the tests of one module. Keep the globs fixed across runs. |
| 189 | `--format` narrows the walk to some languages, for example when the source mixes the code being ported with build scripts. |
| 190 | `--semantic-model` and `--semantic-url` pick another embedding model or an OpenAI-compatible API (`npx jscpd --semantic-models` lists the models with calibrated thresholds). Keep the model fixed across runs, because each model scores on its own scale. |
| 191 | The exit code is 0 whatever the progress, so the command does not fail a build on its own. |
| 192 | |
| 193 | ## Rules |
| 194 | |
| 195 | Keep the two paths, their order, the options and the model the same from run to run, or the numbers stop being comparable. |
| 196 | Keep vendored dependencies, installed packages and build output out of both paths: through a `.gitignore` in a git repository, which you suggest to the user, or with `--ignore`. |
| 197 | Port the tests before the code, and a function only together with the tests the coverage map binds to it. |
| 198 | Treat the report as a map of what to read and what is left. It does not prove the port is correct; tests do. |
| 199 | Never game the percentage: no stubs, no renames made only for jscpd, no deleting source functions to shrink the denominator without the user's decision. |
| 200 | When the report and your reading of the code disagree, trust the code and tell the user which pair looked wrong. |
| 201 |
Discussion
Alternatives
Browse more free Claude skills or everything in Development.