Carbon-13 metabolic flux inference skill

Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics.

by K-Dense-AI·MIT license·★ 45,977 Stars on the repo·GitHub ↗

Use now

Files of Carbon-13 metabolic flux inference

K-Dense-AI/main1 file shown
SKILL.md
Show the full text218 lines

Carbon-13 metabolic flux inference

Turn reviewed carbon maps, explicit tracer mixtures, and corrected labeling measurements into feasible flux estimates and evidence about which fluxes the experiment constrains. Use the bundled solver rather than reconstructing isotope balances or fitting each reaction independently. It runs mfapy's EMU forward simulator and fits fluxes in the mass-balanced feasible space with SciPy. It does not use an FBA objective.

Scope and required evidence

This implementation supports metabolic and isotopic steady state, a single shared flux state across one or more tracer experiments, nonnegative one-way reaction fluxes, and carbon-subset mass distributions. Reversible reactions are two separately mapped directions. Measurement error is Gaussian with a supplied covariance or a disclosed diagonal approximation.

Before fitting, obtain:

  • The carbon network and the source of each atom assignment. Stoichiometry alone does not specify where labeled atoms go. Record compartments as separate metabolite IDs.
  • Evidence for both steady-state assumptions. Stable metabolite abundance does not establish isotopic steady state. Time-course labeling requires INST-MFA with pool sizes and initial labeling; do not average it into this solver.
  • Every carbon input's positional isotopomer distribution, including unlabeled supplements, bicarbonate/CO2 when assimilated, and tracer impurity.
  • Fragment carbon assignments, natural-abundance correction history, and uncertainty of the reported mean. Raw peak intensities, derivatized spectra, and MS/MS transitions require validated preprocessing before these inputs can be constructed.
  • Flux units, extracellular rate measurements or a stated relative-flux reference, and biologically justified bounds. Label fractions alone cannot set an absolute rate.

If necessary information is missing, name it and prepare the input template; do not invent a fragment assignment, atom map, isotope correction, or measurement error. Read references/input-contract.md when preparing inputs. Read references/inference.md before interpreting an actual fit.

Install the tested engine

Run in the user's analysis directory. Set SKILL_DIR to this skill's installed directory, using the actual resolved path. Keep environments and generated results outside the skill.

uv venv --python 3.12 .venv-mfa
uv pip install --python .venv-mfa/bin/python -r "$SKILL_DIR/assets/requirements.txt"

The following commands use .venv-mfa/bin/python; on Windows use the environment's Scripts/python.exe. mfapy is installed from an immutable Git revision because it is not distributed on PyPI. Installation executes dependency build code; model inputs are data, not user-supplied Python. The adapter restricts identifiers and atom-map syntax before they reach mfapy's internally generated numerical functions.

The pinned commit matched upstream master on 2026-09-30. Its README labels the latest change "064", but its installed distribution still reports 0.6.3; retain the Git commit alongside the package version in an analysis record. The refreshed NumPy/SciPy pins require Python 3.12 or later; the commands above use the tested 3.12 environment. See the reviewed forward-model contract in references/inference.md.

Workflow

  1. Prepare explicit inputs. Copy a relevant model asset into the analysis directory, then replace its scientific content only from reviewed evidence. The bundled models are demonstrations, not validated organism-specific reconstructions. Use a separate dataset for each biological condition; jointly fit tracer replicates only when their biological flux state is defensibly shared.

  2. Check the contract and feasibility.

    .venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" check \
      --model model.json --data measurements.json --output input-check.json
    

    This checks atom counts and conservation, fragments, tracer sums, uncertainty matrices, bounds, and steady-state mass-balance feasibility. It cannot verify that a chemically consistent atom map is biologically correct or that a sample reached steady state.

  3. Exercise the forward model. Supply one mass-balanced flux vector in the declared units. Compare predicted labeling with a reference or independently derived limits.

    .venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" simulate \
      --model model.json --data measurements.json --fluxes fluxes.json \
      --output simulated-mdvs.json
    
  4. Fit and profile the fluxes relevant to the question.

    .venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" fit \
      --model model.json --data measurements.json --starts 12 --seed 2026 \
      --profile v3 --profile v7 --profile-points 31 --profile-starts 6 \
      --output fit.json
    

    Replace v3 and v7 with actual reaction IDs. Each profile point fixes that reaction and reoptimizes nuisance fluxes. For nonlinear networks, repeat with a different seed and more starts before interpreting a profile. A small residual is not an identifiability result.

  5. Inspect the evidence. Check failed starts, residual patterns, mass balance, active bounds, local sensitivity rank, and profile status. Report threshold-crossing brackets at their actual grid resolution. Refine the grid if they are too coarse. Each requested profile gives a one-flux interval under the stated error model; multiple 95% profiles are not a simultaneous 95% region for the whole network. If a profile finds a better solution than the baseline, rerun the fit; do not publish the stale intervals. A failed profile point is unknown, not excluded by the data.

  6. Deliver a bounded scientific result. Include model and data hashes, package versions, source/correction provenance, units and reference flux, fitted predictions, residual diagnostics, profile plots or a table, and the unresolved flux combinations. Retain the JSON artifact. Separate point estimates supported by the data from arbitrary optimizer choices along a flat direction. Suggest additional measurements only after testing that their predicted labeling changes along that direction.

Worked examples

These executable examples use synthetic, tracer-only data. There is no hidden natural- abundance correction, and the tracer proportions already include unlabeled material.

Recover a pathway split; then remove the informative measurement

The analytical two-route model sends a two-carbon substrate through either a carbon-preserving or a carbon-swapping route. Uptake is fixed to 100. An 80% carbon-1 labeled feed and a carbon-1 fragment with M+1 = 0.56 determine the preserving route as 70 and the swapping route as 30.

.venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" fit \
  --model "$SKILL_DIR/assets/branch-model.json" \
  --data "$SKILL_DIR/assets/branch-identifiable.json" \
  --profile straight --profile-points 41 --output branch-fit.json

.venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" fit \
  --model "$SKILL_DIR/assets/branch-model.json" \
  --data "$SKILL_DIR/assets/branch-unresolved.json" \
  --profile straight --output unresolved-fit.json

The first fit recovers approximately 70/30. Under its declared Gaussian error model, the analytical 95% interval for straight is about 67.55–72.45; the script reports grid brackets enclosing the threshold crossings. The second fit has only the whole- molecule distribution, which is identical for the two routes. Expect local rank zero and unresolved_within_bounds; its returned split is an arbitrary optimum.

assets/branch-fluxes.json supplies the 70/30 forward-simulation vector.

Reproduce a published cyclic-network calculation
.venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" simulate \
  --model "$SKILL_DIR/assets/tca-model.json" \
  --data "$SKILL_DIR/assets/tca-tracer.json" \
  --fluxes "$SKILL_DIR/assets/tca-fluxes.json" --output tca-simulation.json

.venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" fit \
  --model "$SKILL_DIR/assets/tca-model.json" \
  --data "$SKILL_DIR/assets/tca-reference-mdv.json" \
  --profile v3 --profile v7 --output tca-fit.json

The first command reproduces the published rounded glutamate MDV [0.3464, 0.2695, 0.2708, 0.0807, 0.0286, 0.0039]. The second uses synthetic reference measurements to recover the glutamate branch flux near 50, while recognizing that this labeling does not resolve the fumarate/oxaloacetate exchange. A constraint-induced upper edge is not evidence of a measurement-determined exchange interval.

Interpretation boundaries

  • A positional isotopomer string runs carbon 1 to carbon N from left to right. "100000" means carbon-1 labeled glucose. A mass distribution alone cannot specify that positional mixture. The adapter handles mfapy's reversed integer-bit ordering.
  • Natural-abundance correction and tracer-purity correction are different operations. Inputs must be in the documented tracer-only basis, with tracer impurity represented consistently in source mixtures. Do not correct the same contribution twice.
  • An N-carbon mass distribution has at most N independent components because it sums to one. The tool removes one bin and uses the reduced covariance. Retain cross-bin correlations when available. Diagonal SEM fits are explicitly approximate.
  • The symmetric flag means equal averaging of identity and complete carbon-order reversal, as in the bundled fumarate/succinate map. It is not arbitrary molecular symmetry. Other permutations need an explicitly supported model representation.
  • Unsupported in this CLI: nonstationary MFA, isotope effects on reaction rates, unmodeled pools or compartments, MS/MS joint distributions, multi-element isotope correction, fractional carbon stoichiometry/pseudo-reactions, and organism-scale performance guarantees. For these, use a validated specialized model/engine and retain the same input/provenance and identifiability discipline.

Implementation and validation

scripts/mfa.py is the CLI. scripts/_mfa_model.py validates inputs and adapts them to the mfapy EMU simulator; scripts/_mfa_fit.py handles feasible flux coordinates, multistart optimization, diagnostic rank, and profile calculations. The engine is pinned in assets/requirements.txt.

The repository suite at tests/13c-metabolic-flux/ checks the published reference, analytical split recovery and likelihood profiles, unresolved routes and exchange, omitted-bin invariance with correlated errors, parallel tracers, absolute-rate anchoring, repeated-substrate condensation, symmetry, invalid maps, and CLI behavior. These checks establish the tested numerical behavior, not biological validation of a user's model or a measured advantage over any particular language model.

Sources

1---
2name: 13c-metabolic-flux
3description: Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics. Use for 13C-MFA, carbon tracing, mass isotopomer distributions (MDVs/MIDs), positional isotopomers, parallel tracer experiments, and determining whether labeling data constrain a pathway flux. Distinguishes measured-label inference from COBRA flux balance analysis and flags experiments requiring nonstationary MFA.
4license: MIT
5compatibility: Python 3.12 with uv and Git for installation. Tested with mfapy 0.6.3 at a10433af16682386548b360297e2476152d46ede, NumPy 2.5.3, SciPy 1.18.1, and NLopt 2.11.0. Network access is needed only to install public dependencies. Inference runs locally without credentials; inputs are JSON.
6metadata:
7 version: "1.2"
8 skill-author: K-Dense Inc.
9 last-reviewed: "2026-09-30"
10---
11 
12# Carbon-13 metabolic flux inference
13 
14Turn reviewed carbon maps, explicit tracer mixtures, and corrected labeling measurements
15into feasible flux estimates and evidence about which fluxes the experiment constrains.
16Use the bundled solver rather than reconstructing isotope balances or fitting each
17reaction independently. It runs mfapy's EMU forward simulator and fits fluxes in the
18mass-balanced feasible space with SciPy. It does not use an FBA objective.
19 
20## Scope and required evidence
21 
22This implementation supports **metabolic and isotopic steady state**, a single shared
23flux state across one or more tracer experiments, nonnegative one-way reaction fluxes,
24and carbon-subset mass distributions. Reversible reactions are two separately mapped
25directions. Measurement error is Gaussian with a supplied covariance or a disclosed
26diagonal approximation.
27 
28Before fitting, obtain:
29 
30- The carbon network and the source of each atom assignment. Stoichiometry alone does
31 not specify where labeled atoms go. Record compartments as separate metabolite IDs.
32- Evidence for both steady-state assumptions. Stable metabolite abundance does not
33 establish isotopic steady state. Time-course labeling requires INST-MFA with pool
34 sizes and initial labeling; do not average it into this solver.
35- Every carbon input's positional isotopomer distribution, including unlabeled
36 supplements, bicarbonate/CO2 when assimilated, and tracer impurity.
37- Fragment carbon assignments, natural-abundance correction history, and uncertainty
38 of the **reported mean**. Raw peak intensities, derivatized spectra, and MS/MS
39 transitions require validated preprocessing before these inputs can be constructed.
40- Flux units, extracellular rate measurements or a stated relative-flux reference,
41 and biologically justified bounds. Label fractions alone cannot set an absolute rate.
42 
43If necessary information is missing, name it and prepare the input template; do not
44invent a fragment assignment, atom map, isotope correction, or measurement error.
45Read [references/input-contract.md](references/input-contract.md) when preparing inputs.
46Read [references/inference.md](references/inference.md) before interpreting an actual fit.
47 
48## Install the tested engine
49 
50Run in the user's analysis directory. Set `SKILL_DIR` to this skill's installed directory,
51using the actual resolved path. Keep environments and generated results outside the skill.
52 
53```bash
54uv venv --python 3.12 .venv-mfa
55uv pip install --python .venv-mfa/bin/python -r "$SKILL_DIR/assets/requirements.txt"
56```
57 
58The following commands use `.venv-mfa/bin/python`; on Windows use the environment's
59`Scripts/python.exe`. mfapy is installed from an immutable Git revision because it is
60not distributed on PyPI. Installation executes dependency build code; model inputs
61are data, not user-supplied Python. The adapter restricts identifiers and atom-map
62syntax before they reach mfapy's internally generated numerical functions.
63 
64The pinned commit matched upstream `master` on 2026-09-30. Its README labels the
65latest change "064", but its installed distribution still reports `0.6.3`; retain
66the Git commit alongside the package version in an analysis record. The refreshed
67NumPy/SciPy pins require Python 3.12 or later; the commands above use the tested 3.12
68environment. See the reviewed forward-model contract in
69[references/inference.md](references/inference.md).
70 
71## Workflow
72 
731. **Prepare explicit inputs.** Copy a relevant model asset into the analysis directory,
74 then replace its scientific content only from reviewed evidence. The bundled models
75 are demonstrations, not validated organism-specific reconstructions. Use a separate
76 dataset for each biological condition; jointly fit tracer replicates only when their
77 biological flux state is defensibly shared.
782. **Check the contract and feasibility.**
79 
80 ```bash
81 .venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" check \
82 --model model.json --data measurements.json --output input-check.json
83 ```
84 
85 This checks atom counts and conservation, fragments, tracer sums, uncertainty
86 matrices, bounds, and steady-state mass-balance feasibility. It cannot verify that a
87 chemically consistent atom map is biologically correct or that a sample reached steady state.
883. **Exercise the forward model.** Supply one mass-balanced flux vector in the declared
89 units. Compare predicted labeling with a reference or independently derived limits.
90 
91 ```bash
92 .venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" simulate \
93 --model model.json --data measurements.json --fluxes fluxes.json \
94 --output simulated-mdvs.json
95 ```
96 
974. **Fit and profile the fluxes relevant to the question.**
98 
99 ```bash
100 .venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" fit \
101 --model model.json --data measurements.json --starts 12 --seed 2026 \
102 --profile v3 --profile v7 --profile-points 31 --profile-starts 6 \
103 --output fit.json
104 ```
105 
106 Replace `v3` and `v7` with actual reaction IDs. Each profile point fixes that reaction
107 and reoptimizes nuisance fluxes. For nonlinear networks, repeat with a different seed
108 and more starts before interpreting a profile. A small residual is not an
109 identifiability result.
1105. **Inspect the evidence.** Check failed starts, residual patterns, mass balance,
111 active bounds, local sensitivity rank, and profile status. Report threshold-crossing
112 brackets at their actual grid resolution. Refine the grid if they are too coarse.
113 Each requested profile gives a one-flux interval under the stated error model;
114 multiple 95% profiles are not a simultaneous 95% region for the whole network.
115 If a profile finds a better solution than the baseline, rerun the fit; do not publish
116 the stale intervals. A failed profile point is unknown, not excluded by the data.
1176. **Deliver a bounded scientific result.** Include model and data hashes, package
118 versions, source/correction provenance, units and reference flux, fitted predictions,
119 residual diagnostics, profile plots or a table, and the unresolved flux combinations.
120 Retain the JSON artifact. Separate point estimates supported by the data from arbitrary
121 optimizer choices along a flat direction. Suggest additional measurements only after
122 testing that their predicted labeling changes along that direction.
123 
124## Worked examples
125 
126These executable examples use synthetic, tracer-only data. There is no hidden natural-
127abundance correction, and the tracer proportions already include unlabeled material.
128 
129### Recover a pathway split; then remove the informative measurement
130 
131The analytical two-route model sends a two-carbon substrate through either a
132carbon-preserving or a carbon-swapping route. Uptake is fixed to 100. An 80% carbon-1
133labeled feed and a carbon-1 fragment with M+1 = 0.56 determine the preserving route
134as 70 and the swapping route as 30.
135 
136```bash
137.venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" fit \
138 --model "$SKILL_DIR/assets/branch-model.json" \
139 --data "$SKILL_DIR/assets/branch-identifiable.json" \
140 --profile straight --profile-points 41 --output branch-fit.json
141 
142.venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" fit \
143 --model "$SKILL_DIR/assets/branch-model.json" \
144 --data "$SKILL_DIR/assets/branch-unresolved.json" \
145 --profile straight --output unresolved-fit.json
146```
147 
148The first fit recovers approximately 70/30. Under its declared Gaussian error model,
149the analytical 95% interval for `straight` is about 67.55–72.45; the script reports
150grid brackets enclosing the threshold crossings. The second fit has only the whole-
151molecule distribution, which is identical for the two routes. Expect local rank zero
152and `unresolved_within_bounds`; its returned split is an arbitrary optimum.
153 
154`assets/branch-fluxes.json` supplies the 70/30 forward-simulation vector.
155 
156### Reproduce a published cyclic-network calculation
157 
158```bash
159.venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" simulate \
160 --model "$SKILL_DIR/assets/tca-model.json" \
161 --data "$SKILL_DIR/assets/tca-tracer.json" \
162 --fluxes "$SKILL_DIR/assets/tca-fluxes.json" --output tca-simulation.json
163 
164.venv-mfa/bin/python "$SKILL_DIR/scripts/mfa.py" fit \
165 --model "$SKILL_DIR/assets/tca-model.json" \
166 --data "$SKILL_DIR/assets/tca-reference-mdv.json" \
167 --profile v3 --profile v7 --output tca-fit.json
168```
169 
170The first command reproduces the published rounded glutamate MDV
171`[0.3464, 0.2695, 0.2708, 0.0807, 0.0286, 0.0039]`.
172The second uses synthetic reference measurements to recover the glutamate branch
173flux near 50, while recognizing that this labeling does not resolve the
174fumarate/oxaloacetate exchange. A constraint-induced upper edge is not evidence of
175a measurement-determined exchange interval.
176 
177## Interpretation boundaries
178 
179- A positional isotopomer string runs **carbon 1 to carbon N from left to right**.
180 `"100000"` means carbon-1 labeled glucose. A mass distribution alone cannot specify
181 that positional mixture. The adapter handles mfapy's reversed integer-bit ordering.
182- Natural-abundance correction and tracer-purity correction are different operations.
183 Inputs must be in the documented tracer-only basis, with tracer impurity represented
184 consistently in source mixtures. Do not correct the same contribution twice.
185- An N-carbon mass distribution has at most N independent components because it sums
186 to one. The tool removes one bin and uses the reduced covariance. Retain cross-bin
187 correlations when available. Diagonal SEM fits are explicitly approximate.
188- The `symmetric` flag means equal averaging of identity and **complete carbon-order
189 reversal**, as in the bundled fumarate/succinate map. It is not arbitrary molecular
190 symmetry. Other permutations need an explicitly supported model representation.
191- Unsupported in this CLI: nonstationary MFA, isotope effects on reaction rates,
192 unmodeled pools or compartments, MS/MS joint distributions, multi-element isotope
193 correction, fractional carbon stoichiometry/pseudo-reactions, and organism-scale
194 performance guarantees. For these, use a validated specialized model/engine and
195 retain the same input/provenance and identifiability discipline.
196 
197## Implementation and validation
198 
199`scripts/mfa.py` is the CLI. `scripts/_mfa_model.py` validates inputs and adapts them to
200the mfapy EMU simulator; `scripts/_mfa_fit.py` handles feasible flux coordinates,
201multistart optimization, diagnostic rank, and profile calculations.
202The engine is pinned in [assets/requirements.txt](assets/requirements.txt).
203 
204The repository suite at `tests/13c-metabolic-flux/` checks the published reference,
205analytical split recovery and likelihood profiles, unresolved routes and exchange,
206omitted-bin invariance with correlated errors, parallel tracers, absolute-rate
207anchoring, repeated-substrate condensation, symmetry, invalid maps, and CLI behavior.
208These checks establish the tested numerical behavior, not biological validation of a
209user's model or a measured advantage over any particular language model.
210 
211## Sources
212 
213- [mfapy source at the tested revision](https://github.com/fumiomatsuda/mfapy/tree/a10433af16682386548b360297e2476152d46ede), version 0.6.3; [API documentation](https://fumiomatsuda.github.io/mfapy-document/mfapy.html).
214- Matsuda et al. (2021), [mfapy: An open-source Python package for 13C-based metabolic flux analysis](https://doi.org/10.1016/j.mec.2021.e00177).
215- Antoniewicz, Kelleher, and Stephanopoulos (2007), [Elementary metabolite units (EMU): A novel framework for modeling isotopic distributions](https://doi.org/10.1016/j.ymben.2006.09.001).
216- The TCA network is adapted from mfapy's MIT-licensed example files; the attribution
217 and full notice are in [assets/mfapy-license.txt](assets/mfapy-license.txt).
218 

Discussion

Alternatives