Files
patient_leakage_detection/plan.md
T
rpotter6298 35cbd9ac3c 2026001
2026-07-01 17:35:58 +02:00

117 lines
6.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Implementation Plan — Potter et al. Data Leakage Reproduction
## Manuscript pipeline (from docx)
1. **Feature extraction** — 5 CNNs (VGG16, DenseNet121, EfficientNetB1, MobileNetV2, ResNet50), frozen, include_top=False
2. **PCA + t-SNE** — visualize feature space (Figure 5 / S2)
3. **Patient clustering** — K-means on 64×64 grayscale thumbnails, per-class (15/40/55)
4. **Random Forest** — feature importance ranking (2000 trees → we use 500, equivalent)
5. **SVM with GridSearchCV** — grid over γ per nfeatures, C=10 fixed, 5-fold CV
6. **Image-level vs patient-level** — compare test accuracy with both split types
7. **20-seed repetition** — boxplots showing distribution (Figure 6 / S4)
## Current status
### ✅ Implemented & Working
| Component | Location | Notes |
|---|---|---|
| Feature extraction (5 models, PyTorch) | `classes/features.py` | ResNet50 GAP stripped (100K-d). EfficientNetB1 dim error documented |
| Patient clustering (K-means) | `classes/patient_identifier.py` | Thumbnail + feature-based modes |
| Classification pipeline | `classes/classifier.py` | `PatientLeakageClassifier.run()` — RF→GridSearchCV→SVM |
| t-SNE visualization | `scripts/visualizations/simple_patient_tsne.py` | Per-class patient coloring |
| Figure 6 (20-seed boxplots) | `scripts/visualizations/figure6.py` | Image vs patient distributions |
| Figure S4 (CV accuracy curves) | `scripts/visualizations/figure_s4.py` | Per-model, log₁₀ x-scale, patient-level |
| C×γ grid search | archived | Confirmed C=10 is adequate |
| SVM class-weight analysis | archived | Confirmed no benefit |
| Task06 NIfTI dataset loader | `classes/nifti_dataset.py` | HU windowing, orientation control |
| Clustering validation (Task06) | `scripts/validate_clustering_task06.py` | Thumbnail + feature K-means vs ground truth |
| Clustering comparison (IQ-OTH) | `scripts/analysis/compare_clustering_methods.py` | Thumbnail vs feature vs v8 reference |
| Siamese patient matching | `classes/siamese.py` | `SiamesePatientMatcher` class |
| Siamese training | `scripts/train_siamese.py` | Task06 train/test split, hard negatives |
### 🚧 In Progress
| Component | Status |
|---|---|
| Siamese-based patient manifest for IQ-OTH | `scripts/siamese_identify.py` written, needs trained model |
### ❌ Still Needed
| Manuscript Figure/Table | What we need | Priority |
|---|---|---|
| **Figure 5** (PCA + t-SNE) | Per-model t-SNE plots, matching manuscript style | Medium |
| **Figure S1** (image-level CV curves) | 5 panels, image-level split, log₁₀ x-scale. Nearly identical to S4 but with image split | Low |
| **Figure S2** (PCA/t-SNE per model) | Like Figure 5 but for all 5 models | Medium |
| **Figure S3** (example cluster images) | 5 example slices from a single K-means cluster per class | Low |
| **Figure S5** (confusion matrix) | VGG16 patient-level confusion matrix | Low |
| **Table 2 comparison** | Run classification with thumbnail K-means manifest to match manuscript numbers | High |
| **Siamese results integration** | Once trained: run siamese_identify, compare manifest against K-means manifests | High |
| **Final patient manifest** | Choose best method, produce canonical patient assignments | High |
## Proposed final structure
```
patient_leakage_detector/
├── classes/
│ ├── features.py # FeatureExtractor (5 CNNs)
│ ├── patient_identifier.py # PatientIdentifier (K-means clustering)
│ ├── classifier.py # PatientLeakageClassifier (RF→SVM pipeline)
│ ├── nifti_dataset.py # NiftiSliceDataset (Task06 loader)
│ └── siamese.py # SiamesePatientMatcher (learned matching)
├── scripts/
│ ├── classification.py # Core: single-seed classification
│ ├── train_siamese.py # Core: siamese training on Task06
│ ├── siamese_identify.py # Core: IQ-OTH patient manifest via siamese
│ ├── validate_clustering_task06.py # Core: Task06 validation
│ │
│ ├── visualizations/
│ │ ├── figure5.py # [TODO] PCA + t-SNE plots
│ │ ├── figure6.py # ✅ 20-seed boxplots
│ │ ├── figure_s1.py # [TODO] Image-level CV curves
│ │ ├── figure_s2.py # [TODO] PCA/t-SNE per model
│ │ ├── figure_s3.py # [TODO] Example cluster images
│ │ ├── figure_s4.py # ✅ Patient-level CV curves
│ │ ├── figure_s5.py # [TODO] Confusion matrix
│ │ └── simple_patient_tsne.py # ✅ t-SNE with patient coloring
│ │
│ └── analysis/
│ ├── compare_clustering_methods.py # ✅ Thumbnail vs feature vs v8
│ └── verify_feature_dims.py # ✅ TF dimension verification
├── features/
│ ├── VGG16_features.npz # Extracted features (all 5 models)
│ └── task06_lung/ # Cached Task06 features
├── models/
│ └── siamese_resnet18.pt # Trained siamese model
├── results/
│ ├── simple_patient_manifest.csv # Feature-based K-means manifest
│ ├── siamese_manifest.csv # [TODO] Siamese-based manifest
│ ├── classification_results.json # Single-seed results (all 5 models)
│ ├── figure6_data.json # 20-seed data
│ ├── task06_clustering_validation.json # Task06 validation results
│ └── ...
├── plots/
│ ├── figure6.png
│ ├── figure_s4.png
│ └── ...
├── plan.md # This file
├── notes.md # Review notes for supervisor
├── .gitignore
└── .archive/ # Superseded scripts and results
```
## Immediate next steps
1. **Finish siamese training** → run `python scripts/siamese_identify.py` → produce `siamese_manifest.csv`
2. **Run classification with thumbnail K-means manifest** to reproduce manuscript Table 2 numbers
3. **Compare manifests**: K-means thumbnail vs K-means feature vs siamese — which gives the best patient-level split?
4. **Choose canonical manifest** and produce final classification results
5. **Generate remaining figures** (Figure 5, S1, S2, S3, S5)
6. **Draft findings for supervisor discussion**