CellProfiler quantitative microscopy skill

Runs reproducible CellProfiler microscopy pipelines for nuclear segmentation, cell counts, and per-object fluorescence measurements.

by K-Dense-AI·MIT license·★ 45,977 Stars on the repo·GitHub ↗

Use now

Files of CellProfiler quantitative microscopy

K-Dense-AI/main1 file shown
SKILL.md
Show the full text115 lines

CellProfiler quantitative microscopy

Use this skill when a user needs a repeatable CellProfiler .cppipe, nuclear counts, nuclear fluorescence, or batch microscopy measurements. The bundled assay accepts one 2D grayscale TIFF nuclear channel per field, with black-is-zero (MINISBLACK) pixels and bright nuclei on a dark background. Palette and white-is-zero TIFFs need an explicit conversion. For volumetric segmentation, multichannel cell painting, or tissue-specific models, design a separate pipeline and validate those assumptions rather than silently projecting or splitting the images.

The official application and manual remain 4.2.8. PyPI publishes 4.2.8.1; its seven modules used here and embedded Threshold module match the 4.2.8 source, but this review did not execute that native distribution. Keep the helper environment separate from CellProfiler's older dependency stack; see the runtime reference for the verification boundary.

Workflow

  1. Establish the acquisition unit: plate, well, site, time point if present, pixel size, nuclear channel identity, camera bit depth, exposure, and biological replicate. Keep original image intensities. Convert proprietary formats explicitly with Bio-Formats before using this helper.
  2. Create the CSV manifest below. image_path is absolute or relative to the manifest; sample IDs use letters, digits, dots, dashes, or underscores; sample IDs and plate/well/site combinations are unique. Use a nonnumeric sample ID such as sample_001: LoadData infers column types and can otherwise turn 001 into 1. Avoid surrounding whitespace in identifiers. TIFFs must be uint8 or uint16, single plane/series/resolution, and nonconstant. The helper rejects RGB, z-stacks, and float images rather than guessing channels.
  3. Use assets/nuclei.cppipe as a starting pipeline: LoadData → IdentifyPrimaryObjects → intensity/size measurements → outline overlay → CSV export. The initial diameter range is 8–80 pixels, with global Otsu thresholding, no threshold smoothing, and border objects excluded. Calibrate this range from representative images and acquisition pixel size before comparing conditions.
  4. Run a small pilot spanning controls, low/high density, dim images, and plate edges. Inspect saved overlays for missed nuclei, splits, merges, and edge exclusions. Adjust thresholding and declumping in CellProfiler, export the tuned .cppipe, and pass --pipeline to preserve it. Do not choose settings separately for each treatment to make their counts agree.
  5. Freeze the tuned pipeline and analyze the batch. Review input saturation warnings, zero counts, count/area distributions, and control behavior. Aggregation for inference belongs at the biological replicate level; thousands of cells from one well are not independent wells.

Run the bounded assay

From this skill directory, create images.csv:

sample_id,image_path,plate,well,site
control_A01_1,images/control_A01_1_DAPI.tif,Plate1,A01,1
python scripts/nuclei_assay.py prepare images.csv load_data.csv
python scripts/nuclei_assay.py run images.csv results --executable cellprofiler
python scripts/nuclei_assay.py summarize results

run requires a fresh/empty output directory and executes CellProfiler with -c -r, a saved pipeline copy, --data-file, output folder, and --done-file. Success requires exit code zero, a Complete marker, and valid measurement tables. It records the command, pipeline checksum, input image checksums, and sample QC in assay_qc.json before execution, retaining failed status and the error if execution or output validation fails. CellProfiler output goes to cellprofiler.log. Rerun in a new output folder. summarize checks CSV contents independently; it does not prove an engine run completed.

Custom pipelines must preserve DNA, Nuclei, Metadata_Sample, integer-dtype scaling, and the unprefixed single-object Image.csv/Nuclei.csv export contract. Keep the required intensity/area measurements. A renamed object set or different intensity scale needs a corresponding helper adaptation, not an unchecked --pipeline substitution.

The executable can also be the CellProfiler application launcher or a local container launcher; see references/runtime-and-qc.md for the container target, filesystem mapping, and verification evidence. prepare and summarize work without CellProfiler.

Interpret the outputs

  • Image.csv: one image/field row, including Count_Nuclei and acquisition metadata.
  • Nuclei.csv: one accepted object per row, with mean/integrated DNA intensity, area, and shape.
  • *_nuclei.png: green nuclear boundaries over the input image for visual QC.
  • pipeline.cppipe and cellprofiler.done: the exact pipeline copy and engine completion marker.
  • assay_qc.json: run status, unique image/object keys, exact counts, finite mean/integrated intensity and positive area checks, field mean area in pixels, and storage saturation flags.

LoadData ignores camera metadata for scaling in this asset and divides by the integer storage maximum: uint8 → 255, uint16 → 65535. A 12-bit camera stored in uint16 therefore has a maximum near 0.0625. Do not compare intensities across different bit depths, exposures, gains, or staining batches without an explicit calibration. A saturated image can pass segmentation while its intensity measurement is unusable. Illumination correction and background subtraction are assay-specific additions; this starter does neither.

The saturation fraction only counts pixels at the storage maximum. A 12-bit detector may saturate at 4095 while the uint16 storage maximum is 65535; inspect the known acquisition ceiling separately. Integrated intensity sums pixel values and may exceed 1; only per-pixel mean intensity is constrained to 0–1. The field's mean nuclear intensity weights each nucleus equally, rather than weighting each pixel equally.

A count check cannot prove correct segmentation. Inspect overlays and independently annotated fields; report boundary exclusions and segmentation errors alongside the biological result. The optional synthetic engine test targets a known three-nucleus example, not assay performance on unseen cell types. It was skipped in the current review because no engine was configured.

Sources

1---
2name: cellprofiler
3description: Runs reproducible CellProfiler microscopy pipelines for nuclear segmentation, cell counts, and per-object fluorescence measurements. Supports image/channel manifests, headless batch execution, segmentation overlays, and measurement QC for 2D fluorescence assays.
4license: MIT
5compatibility: Python 3.12+ with numpy and tifffile for current helper-only dependencies; a separate CellProfiler 4.2.8 application/container for segmentation. Full CellProfiler has older native dependencies. Network access is needed for installation only. No credentials required.
6metadata:
7 version: "1.1"
8 skill-author: K-Dense Inc.
9 upstream-version: "4.2.8"
10 last-reviewed: "2026-09-30"
11---
12 
13# CellProfiler quantitative microscopy
14 
15Use this skill when a user needs a repeatable CellProfiler `.cppipe`, nuclear counts, nuclear
16fluorescence, or batch microscopy measurements. The bundled assay accepts **one 2D grayscale
17TIFF nuclear channel per field**, with black-is-zero (MINISBLACK) pixels and bright nuclei on a
18dark background. Palette and white-is-zero TIFFs need an explicit conversion. For volumetric
19segmentation, multichannel cell painting, or tissue-specific models, design a separate pipeline
20and validate those assumptions rather than silently projecting or splitting the images.
21 
22The official application and manual remain **4.2.8**. PyPI publishes **4.2.8.1**; its seven
23modules used here and embedded Threshold module match the 4.2.8 source, but this review did
24not execute that native distribution. Keep the helper environment separate from CellProfiler's
25older dependency stack; see the runtime reference for the verification boundary.
26 
27## Workflow
28 
291. Establish the acquisition unit: plate, well, site, time point if present, pixel size, nuclear
30 channel identity, camera bit depth, exposure, and biological replicate. Keep original image
31 intensities. Convert proprietary formats explicitly with Bio-Formats before using this helper.
322. Create the CSV manifest below. `image_path` is absolute or relative to the manifest; sample IDs use letters, digits, dots, dashes, or underscores; sample IDs
33 and plate/well/site combinations are unique. Use a nonnumeric sample ID such as `sample_001`:
34 LoadData infers column types and can otherwise turn `001` into `1`. Avoid surrounding
35 whitespace in identifiers. TIFFs must be uint8 or uint16, single plane/series/resolution, and
36 nonconstant. The helper rejects RGB, z-stacks, and float images rather than guessing channels.
373. Use [assets/nuclei.cppipe](assets/nuclei.cppipe) as a starting pipeline: LoadData →
38 IdentifyPrimaryObjects → intensity/size measurements → outline overlay → CSV export.
39 The initial diameter range is 8–80 **pixels**, with global Otsu thresholding, no threshold
40 smoothing, and border objects excluded. Calibrate this
41 range from representative images and acquisition pixel size before comparing conditions.
424. Run a small pilot spanning controls, low/high density, dim images, and plate edges. Inspect
43 saved overlays for missed nuclei, splits, merges, and edge exclusions. Adjust thresholding
44 and declumping in CellProfiler, export the tuned `.cppipe`, and pass `--pipeline` to preserve
45 it. Do not choose settings separately for each treatment to make their counts agree.
465. Freeze the tuned pipeline and analyze the batch. Review input saturation warnings, zero
47 counts, count/area distributions, and control behavior. Aggregation for inference belongs at
48 the biological replicate level; thousands of cells from one well are not independent wells.
49 
50## Run the bounded assay
51 
52From this skill directory, create `images.csv`:
53 
54```csv
55sample_id,image_path,plate,well,site
56control_A01_1,images/control_A01_1_DAPI.tif,Plate1,A01,1
57```
58 
59```bash
60python scripts/nuclei_assay.py prepare images.csv load_data.csv
61python scripts/nuclei_assay.py run images.csv results --executable cellprofiler
62python scripts/nuclei_assay.py summarize results
63```
64 
65`run` requires a fresh/empty output directory and executes CellProfiler with `-c -r`, a saved
66pipeline copy, `--data-file`, output folder, and `--done-file`. Success requires exit code zero,
67a `Complete` marker, and valid measurement tables. It records the command, pipeline checksum,
68input image checksums, and sample QC in `assay_qc.json` before execution, retaining `failed`
69status and the error if execution or output validation fails. CellProfiler output goes to
70`cellprofiler.log`. Rerun in a new output folder. `summarize` checks CSV contents independently;
71it does not prove an engine run completed.
72 
73Custom pipelines must preserve `DNA`, `Nuclei`, `Metadata_Sample`, integer-dtype scaling, and
74the unprefixed single-object `Image.csv`/`Nuclei.csv` export contract. Keep the required
75intensity/area measurements. A renamed object set or different intensity scale needs a
76corresponding helper adaptation, not an unchecked `--pipeline` substitution.
77 
78The executable can also be the CellProfiler application launcher or a local container launcher;
79see [references/runtime-and-qc.md](references/runtime-and-qc.md) for the container target,
80filesystem mapping, and verification evidence. `prepare` and `summarize` work without CellProfiler.
81 
82## Interpret the outputs
83 
84- `Image.csv`: one image/field row, including `Count_Nuclei` and acquisition metadata.
85- `Nuclei.csv`: one accepted object per row, with mean/integrated DNA intensity, area, and shape.
86- `*_nuclei.png`: green nuclear boundaries over the input image for visual QC.
87- `pipeline.cppipe` and `cellprofiler.done`: the exact pipeline copy and engine completion marker.
88- `assay_qc.json`: run status, unique image/object keys, exact counts, finite mean/integrated
89 intensity and positive area checks, field mean area in pixels, and storage saturation flags.
90 
91LoadData ignores camera metadata for scaling in this asset and divides by the integer storage
92maximum: uint8 → 255, uint16 → 65535. A 12-bit camera stored in uint16 therefore has a maximum
93near 0.0625. Do not compare intensities across different bit depths, exposures, gains, or staining
94batches without an explicit calibration. A saturated image can pass segmentation while its
95intensity measurement is unusable. Illumination correction and background subtraction are
96assay-specific additions; this starter does neither.
97 
98The saturation fraction only counts pixels at the **storage maximum**. A 12-bit detector may
99saturate at 4095 while the uint16 storage maximum is 65535; inspect the known acquisition ceiling
100separately. Integrated intensity sums pixel values and may exceed 1; only per-pixel mean
101intensity is constrained to 0–1. The field's mean nuclear intensity weights each nucleus
102equally, rather than weighting each pixel equally.
103 
104A count check cannot prove correct segmentation. Inspect overlays and independently annotated
105fields; report boundary exclusions and segmentation errors alongside the biological result.
106The optional synthetic engine test targets a known three-nucleus example, not assay performance
107on unseen cell types. It was skipped in the current review because no engine was configured.
108 
109## Sources
110 
111- [Official example pipelines](https://cellprofiler.org/examples): choose an assay-specific starting point.
112- [CellProfiler 4.2.8 manual](https://cellprofiler-manual.s3.amazonaws.com/CellProfiler-4.2.8/index.html): module settings and interpretation.
113- [Headless batch processing](https://cellprofiler-manual.s3.amazonaws.com/CellProfiler-4.2.8/help/other_batch.html): command-line execution.
114- [Current application download](https://cellprofiler.org/releases) and [PyPI distribution](https://pypi.org/project/cellprofiler/): distinct release targets.
115 

Discussion

Alternatives

Aeon Time Series Machine LearningThis skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.Science · MITdeepTools: NGS Data Analysis ToolkitNGS analysis toolkit. BAM to bigWig conversion, QC (correlation, PCA, fingerprints), heatmaps/profiles (TSS, peaks), for ChIP-seq, RNA-seq, ATAC-seq visualization.Science · MITExploratory data analysisPerform bounded, local exploratory analysis of explicitly supported scientific files. Use for redacted CSV/TSV/JSON profiles; optional NumPy, HDF5, FASTA/FASTQ, and basic image metadata inspection; missingness/leakage audits; outlier and transformation sensitivity; and rigorous EDA report scaffolds. Other domain formats are reference-only and unknown formats fail closed.Science · MITNeuropixels Data AnalysisAnalyze Neuropixels extracellular recordings end-to-end with SpikeInterface. Covers loading SpikeGLX/Open Ephys/NWB data, preprocessing, drift/motion correction, Kilosort4 (and CPU) spike sorting, quality metrics, and unit curation (threshold-based, model-based UnitRefine, and AI-assisted visual review). Use when working with Neuropixels 1.0/2.0 recordings, spike sorting, or extracellular electrophysiology analysis.Science · MIT