The Broadest-Coverage Unified PyTorch Library for 3D Point Cloud Deep Learning
Classification · Semantic Segmentation · Part Segmentation · Self-Supervised Pre-training · Parameter-Efficient Fine-Tuning · Few-Shot
Paper: LIDARLearn: A Unified Deep Learning Library for 3D Point Cloud Classification, Segmentation, and Self-Supervised Representation Learning · Demo: Google Drive walkthrough
📄 See what LIDARLearn produces → docs/model_comparison_tables.pdf — a real, publication-ready PDF auto-generated by the library from raw training logs. No manual typesetting. This one-click report pipeline is LIDARLearn's headline contribution.
LIDARLearn ships the most complete off-the-shelf catalogue of 3D point cloud deep-learning methods we're aware of — 56 model configurations covering 29 supervised backbones, 7 self-supervised pre-training methods, and 5 parameter-efficient fine-tuning (PEFT) strategies (DAPT, IDPT, PPT, GST, VPT-Deep) composable with 4 of the SSL backbones (ACT, PointGPT, Point-MAE, ReCon) — all exposed through a single registry-based framework with YAML-driven configs, standardised runners, automatic LaTeX/CSV reporting, and a 2 200+ test pytest suite.
Designed for researchers and engineers who need to run apples-to-apples benchmarks across heterogeneous 3D point-cloud sources in the same pipeline:
- General 3D data — ModelNet40, ShapeNet-55/34, ShapeNet Part
- Indoor 3D scans — S3DIS (Matterport-style RGB + XYZ, semantic segmentation)
- Terrestrial Laser Scanning (TLS) — STPCTLS (tree species classification on high-density ground-based scans)
- Aerial / helicopter LiDAR (ALS) — HELIALS (tree species from sparse-to-dense aerial point clouds)
Same YAML schema, same CLI, same report generators for all of them — swap the dataset, keep the model and training recipe constant. Supported tasks: classification, semantic segmentation, part segmentation, and few-shot classification, with built-in stratified K-fold cross-validation and Friedman/Nemenyi statistical analysis.
The matrix below enumerates every (model × fine-tuning strategy) pair shipped with ready-to-run YAML configs. Columns: Cls — object classification (ModelNet40 / STPCTLS / HELIALS / ModelNetFewShot), PartSeg — part segmentation (ShapeNetParts), SemSeg — semantic segmentation (S3DIS). "Strategy" is the PEFT adapter applied on top of an SSL backbone: FF = full finetuning (no PEFT), plus 5 PEFT methods — DAPT, IDPT, PPT, GST (PointGST), VPT-Deep. The 7 SSL backbones also support pre-training on ShapeNet-55.
| Category | Model | Strategy | Cls | PartSeg | SemSeg |
|---|---|---|---|---|---|
| Point-based | PointNet | — | ✓ | ✓ | ✓ |
| PointNet2-SSG | — | ✓ | ✓ | ✓ | |
| PointNet2-MSG | — | ✓ | ✓ | ✓ | |
| SONet | — | ✓ | — | — | |
| PPFNet | — | ✓ | — | — | |
| PointCNN | — | ✓ | — | — | |
| PointWeb | — | ✓ | ✓ | ✓ | |
| PointConv | — | ✓ | ✓ | ✓ | |
| RSCNN | — | ✓ | ✓ | ✓ | |
| PointMLP | — | ✓ | ✓ | ✓ | |
| PointSCNet | — | ✓ | ✓ | ✓ | |
| RepSurf | — | ✓ | ✓ | ✓ | |
| PointKAN | — | ✓ | ✓ | ✓ | |
| DELA | — | ✓ | ✓ | ✓ | |
| RandLA-Net | — | — | ✓ | ✓ | |
| Attention-based | PCT | — | ✓ | ✓ | ✓ |
| P2P | — | ✓ | ✓ | ✓ | |
| PointTNT | — | ✓ | ✓ | ✓ | |
| GlobalTransformer | — | ✓ | ✓ | ✓ | |
| PVT | — | ✓ | ✓ | ✓ | |
| PointTransformer | — | ✓ | ✓ | ✓ | |
| PointTransformerV2 | — | ✓ | ✓ | ✓ | |
| PointTransformerV3 | — | ✓ | ✓ | ✓ | |
| Graph-based | DGCNN | — | ✓ | ✓ | ✓ |
| DeepGCN | — | ✓ | ✓ | ✓ | |
| CurveNet | — | ✓ | ✓ | ✓ | |
| GDANet | — | ✓ | ✓ | ✓ | |
| MS-DGCNN | — | ✓ | ✓ | ✓ | |
| KAN-DGCNN | — | ✓ | ✓ | ✓ | |
| MS-DGCNN++ | — | ✓ | ✓ | ✓ | |
| Self-supervised | Point-MAE | FF | ✓ | ✓ | ✓ |
| DAPT | ✓ | ✓ | ✓ | ||
| IDPT | ✓ | ✓ | ✓ | ||
| PPT | ✓ | ✓ | ✓ | ||
| GST | ✓ | ✓ | ✓ | ||
| VPT-Deep | ✓ | — | — | ||
| ACT | FF | ✓ | ✓ | ✓ | |
| DAPT | ✓ | ✓ | ✓ | ||
| IDPT | ✓ | ✓ | ✓ | ||
| PPT | ✓ | ✓ | ✓ | ||
| GST | ✓ | ✓ | ✓ | ||
| VPT-Deep | ✓ | — | — | ||
| ReCon | FF | ✓ | ✓ | ✓ | |
| DAPT | ✓ | ✓ | ✓ | ||
| IDPT | ✓ | ✓ | ✓ | ||
| PPT | ✓ | ✓ | ✓ | ||
| GST | ✓ | ✓ | ✓ | ||
| VPT-Deep | ✓ | — | — | ||
| PointGPT | FF | ✓ | ✓ | ✓ | |
| DAPT | ✓ | — | — | ||
| IDPT | ✓ | — | — | ||
| PPT | ✓ | — | — | ||
| GST | ✓ | — | — | ||
| VPT-Deep | ✓ | — | — | ||
| Point-M2AE | FF | ✓ | ✓ | ✓ | |
| Point-BERT | FF | ✓ | ✓ | ✓ | |
| PCP-MAE | FF | ✓ | ✓ | ✓ |
Total: 56 configurations (29 supervised × FF, 7 SSL × FF, 4 SSL × 5 PEFT = 20). See THIRD_PARTY_NOTICES.md for per-model licences and upstream repositories, and docs/point_cloud_methods.csv for paper titles and venues.
- Unified config system — one YAML per experiment,
_base_inheritance, identical CLI across every model and task. - Cross-validation — stratified K-fold with aggregated metrics (
--run_all_folds). - Friedman / Nemenyi — non-parametric multi-model statistical testing with critical-difference diagrams (
scripts/reports/friedman_significance_report.py). - Automated reporting — publication-ready, standalone LaTeX documents auto-generated from experiment folders (
scripts/reports/). Best value per metric column is auto-bolded (\textbf{}). Every report supports--use_citation_in_tablesto inline\citep{}in the tables, or a single citation paragraph above them by default. - SSL pre-training — ShapeNet-55 pre-training recipes for all 7 SSL backbones.
- PEFT — 5 fine-tuning strategies composable with any SSL backbone via a single
finetuning_strategyfield. - Testing — L1 config-parse · L2 forward-shape · L3 pretrained-load · L4 script-coverage.
- Visualisation — interactive 3D HTML for classification, part-seg, and semantic-seg predictions.
- Confusion matrices — per-run confusion matrix PNG + CSV auto-exported for every classification experiment, so per-class errors are visible without extra tooling.
- Training curves — train/validation loss and accuracy history plotted and saved as PNG after every run, giving an at-a-glance convergence check.
LIDARLearn is validated on Python 3.11 + PyTorch 2.4.1 + CUDA 11.8. The steps below reproduce that exact environment. torch-scatter and torch-cluster are installed separately from the official PyG wheel index because their wheels are pinned to a specific (torch × CUDA) pair — installing them via plain pip install will either fail or pull the wrong build.
conda create -n lidarlearn python=3.11 -y
conda activate lidarlearnpip install torch==2.4.1 torchvision==0.19.1 --index-url https://download.pytorch.org/whl/cu118Pass the torch+CUDA tag in the URL — this is what avoids the "wrong ABI / no CUDA" failures:
pip install torch-scatter torch-cluster -f https://data.pyg.org/whl/torch-2.4.1+cu118.html
pip install torch-geometricVerify they resolved to the CUDA wheels (not CPU-only fallbacks):
python -c "import torch_scatter, torch_cluster; print(torch_scatter.__version__, torch_cluster.__version__)"git clone https://github.com/said-ohamouddou/LIDARLearn.git
cd LIDARLearn
pip install -r requirements.txt(The torch-scatter / torch-cluster / torch-geometric lines in requirements.txt are kept as a safety net and will be a no-op since step 3 already installed them.)
bash extensions/install_extensions.shThis builds 7 extensions: pointnet2_ops, chamfer_dist, dela_cutils, pointops, ptv_modules, clip, and index_max (needed by SO-Net). If you're on CPU-only or don't need the models that depend on them, individual extensions can fail without blocking the rest of the library — the affected models are guarded by _safe_import in models/__init__.py and simply won't register.
PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 pytest tests/test_configs.py -vAll 1100+ config-build tests should pass. If SONet tests fail with Model 'SONet' is not registered, re-run step 5 and check that index_max built successfully.
Tested setup. LIDARLearn has been developed and validated on:
- OS Ubuntu 22.04
- CPU AMD Ryzen 7 5700X (8 cores / 16 threads)
- GPU NVIDIA GeForce RTX 4060 Ti (single-GPU; all runs validated on a single-GPU configuration)
- Python 3.11.10
- PyTorch 2.4.1 (
+cu118) / torchvision 0.19.1 (+cu118) - CUDA 11.8 (nvcc 11.8.89)
- torch-scatter / torch-cluster from
https://data.pyg.org/whl/torch-2.4.1+cu118.html - timm 0.4.5
- ninja (for C++/CUDA extension builds)
Other combinations (Python ≥ 3.8, PyTorch ≥ 1.13, CUDA ≥ 11.3) are likely to work but have not been exhaustively tested. If you use a different torch/CUDA pair, update the URL in step 3 to match — e.g.
torch-2.3.0+cu121for torch 2.3 on CUDA 12.1. Feedback and compatibility reports from other versions are very welcome — please open an issue with your environment details.Multi-GPU / distributed. The codebase wires through
torch.distributed(seeutils/dist_utils.py) but every shipped benchmark was run on a single GPU. Feedback from users running multi-GPU or distributed setups is especially welcome — please open an issue with your launcher, world size, and any adjustments needed.
The 7 self-supervised backbones (Point-MAE, Point-BERT, PointGPT, ACT, ReCon, PCP-MAE, Point-M2AE) need a pretrained checkpoint to fine-tune from. All checkpoints are published together on Google Drive:
Download: Google Drive — LIDARLearn pretrained weights
Place every downloaded .pth into pretrained/ at the repo root:
pretrained/
├── pretrained_mae.pth # Point-MAE — used by cfgs/classification/PointMAE/**
├── pretrained_bert.pth # Point-BERT — cfgs/classification/PointBERT/**
├── pretrained_gpt.pth # PointGPT — cfgs/classification/PointGPT/**
├── pretrained_act.pth # ACT — cfgs/classification/ACT/**
├── pretrained_recon.pth # ReCon — cfgs/classification/RECON/**
├── pretrained_pcp.pth # PCP-MAE — cfgs/classification/PCPMAE/**
├── pretrained_m2ae.pth # Point-M2AE — cfgs/classification/PointM2AE/**
├── act_dvae.pth # ACT dVAE tokeniser (needed only for ACT pre-training)
└── dVAE.pth # Point-BERT dVAE tokeniser (needed only for Point-BERT pre-training)
Pass the relevant checkpoint via --ckpts pretrained/<name>.pth when fine-tuning, e.g.:
python main.py --config cfgs/classification/PointMAE/STPCTLS/stpctls_cv_dapt.yaml \
--ckpts pretrained/pretrained_mae.pth --run_all_folds --seed 42Each loader strips the appropriate state-dict prefix (MAE_encoder., ACT_encoder., transformer_q., GPT_Transformer., …) so checkpoints trained with different SSL objectives all map cleanly onto the finetune model. See tests/test_pretrained_load.py for the verification logic that asserts every loader actually consumes its checkpoint (guards against silent prefix-mismatch failures).
You can skip this section entirely if you only plan to train the 29 supervised backbones — they don't need any pretrained weights.
| Dataset | Task | Loader |
|---|---|---|
| ModelNet40 | Classification (40 classes) | datasets/ModelNetDataset.py |
| ModelNet40-FewShot | Few-shot classification (5/10-way, 10/20-shot) | datasets/ModelNetDatasetFewShot.py |
| ShapeNet-55 | SSL pre-training | datasets/ShapeNet55DatasetPretrain.py |
| ShapeNet Part | Part segmentation (50 classes, 16 categories) | datasets/ShapeNet55Dataset.py |
| S3DIS | Indoor semantic segmentation (13 classes) | datasets/S3DISDataset.py |
| STPCTLS (Göttingen Research Online) | LiDAR tree species (7 classes, TLS point clouds) | datasets/TreeSpeciesDataset.py |
| STPCTLS-CV | STPCTLS with stratified K-fold CV | datasets/TreeSpeciesDatasetCV.py |
| HELIALS (Zenodo 17077256) | Helicopter ALS tree species | datasets/TreeSpeciesDatasetHELIALS.py |
Dataset download and preparation instructions are in DATASET.md. For convenience, the preprocessed STPCTLS data is already included in the repository (under data/) so you can reproduce the tree-species benchmark out of the box without any external download.
# 1. Train PointNet on STPCTLS (simple 80/20 split)
python main.py --config cfgs/classification/PointNet/STPCTLS/stpctls.yaml \
--mode finetune --seed 42 --exp_name pointnet_stpctls
# 2. Train DGCNN on STPCTLS with 5-fold cross-validation
python main.py --config cfgs/classification/DGCNN/STPCTLS/stpctls_cv.yaml \
--mode finetune --seed 42 --run_all_folds --exp_name dgcnn_stpctls_cv
# 3. Pre-train Point-MAE on ShapeNet-55
python main.py --config cfgs/classification/PointMAE/ShapeNet55/shapenet55_pretrain.yaml \
--mode pretrain --exp_name mae_pretrain
# 4. Fine-tune with DAPT on STPCTLS (5-fold CV)
python main.py --config cfgs/classification/PointMAE/STPCTLS/stpctls_cv_dapt.yaml \
--ckpts pretrained/pretrained_mae.pth --run_all_folds --seed 42
# 5. Generate publication-ready LaTeX tables (outputs land in <exp_dir>/latex/)
python scripts/reports/classification_comparison_report.py --exp_dir experiments/STPCTLS
# 6. Friedman / Nemenyi statistical analysis (outputs land in <exp_dir>/friedman/)
python scripts/reports/friedman_significance_report.py --exp_dir experiments/STPCTLSSweep the entire benchmark with a single command:
bash scripts/train_stpctls.sh # all 56 classification configs
bash scripts/train_fewshot.sh # few-shot evaluation
bash scripts/train_s3dis.sh # semantic segmentation
bash scripts/train_shapenetparts.sh # part segmentationTraining a subset of models. The
scripts/train_*.shfiles are plain bash with onepython main.py ...line per model (each preceded by anecho "[N/M] <Name>"header). To run only a subset, open the script and either comment out (#) the models you want to skip, or copy thepython main.py ...line(s) you want directly to your shell. The report scripts pick up whichevercv_summary.csvfiles exist underexperiments/<NAME>/, so partial runs produce valid tables for the models you did train.
Classification / fine-tuning goes through two wired-up runners, exposed in main.py as:
from tools import finetune_run_net as finetune # default
from tools.runner_finetune_test import run_net as finetune_test # 3-way splitSelect with the --runner flag (default runner_finetune):
| Runner | Data splits used | "Best checkpoint" chosen on | Final metric reported on | Use when |
|---|---|---|---|---|
runner_finetune (default) |
dataset.train + dataset.val |
validation | validation | You have a train/val 2-way split (ModelNet40, STPCTLS CV folds, fewshot episodes). Validation is both the selection signal AND the reported number. |
runner_finetune_test |
dataset.train + dataset.val + dataset.test |
validation | test | You have a genuine 3-way split and want the best checkpoint selected on val but the final headline number computed on a held-out test set (cleanest protocol for paper-style reporting). |
# 2-way (default): val metrics are the final numbers
python main.py --config cfgs/classification/DGCNN/STPCTLS/stpctls.yaml --mode finetune --exp_name dgcnn
# 3-way: best ckpt on val, final metrics on test
python main.py --config cfgs/classification/PointMAE/STPCTLS/stpctls_test.yaml \
--mode finetune --runner runner_finetune_test --exp_name pointmae_testBoth runners write the same cv_summary.csv schema so the LaTeX reporters in scripts/reports/ work unchanged. Single-run (non-CV) experiments still produce a cv_summary.csv — the per-metric strings are just formatted as value ± 0.00 (zero std), so downstream tables and Friedman/Nemenyi scripts see a consistent shape whether you ran 1 fold or K folds.
Point-cloud augmentations are implemented in datasets/augmentation.py and applied online — on the GPU, per batch, inside the training loop — not baked into disk files. That means:
- Every epoch sees freshly-sampled augmentations (different rotations, jitter noise, dropout masks) even though the underlying
.xyz/.h5data on disk is fixed. - Applied to training only — validation and test passes skip augmentation (only the always-on unit-sphere normalisation runs there). Guarded at the call site in
tools/runner_finetune.py:255, so every model/config gets the same guarantee without extra code. - GPU-side — transforms take and return a
(B, C, N)or(B, N, C)torch.Tensoralready on CUDA, so there's no CPU→GPU round-trip in the data path.
Enable via --augmentation <name> (default: none). Choices are defined in utils/parser.py and map 1:1 to classes in the augmentation module:
| Flag | What it does |
|---|---|
none |
Identity — no augmentation (default) |
rotate |
Random 3D rotation around all three axes |
scale_translate |
Combined random uniform scaling + translation |
jitter |
Additive Gaussian noise on every point |
scale |
Random uniform scaling only |
translate |
Random translation only |
dropout |
Random point dropout (mask out a fraction of points) |
flip |
Random axis flipping (x or y) |
z_rotate_tree |
Rotation around Z-axis only — right default for upright tree scans (TLS / ALS) where up is meaningful |
Example — full-sweep smoke test with Z-axis rotation for tree data:
python main.py --config cfgs/classification/PointMAE/STPCTLS/stpctls_cv.yaml \
--augmentation z_rotate_tree --run_all_folds --seed 42 \
--exp_name pointmae_stpctls_cv_zrotComposing multiple transforms. The CLI takes a single name, but the underlying get_train_transforms() API accepts a list — compose chains like ['z_rotate_tree', 'jitter', 'scale'] by calling it from a custom training script or by editing the runner. Every composed transform is re-randomised per batch.
Reproducibility. Augmentations use the global torch/numpy RNG, so passing --seed 42 makes a full training run bitwise-reproducible across machines (assuming the same CUDA kernels). The selected transform is logged at the start of training via fmt.print_augmentation(...) so you can confirm from the log what ran.
All three viz scripts live in scripts/visulization/ and emit self-contained interactive HTML files (Plotly-based 3D scatter) plus .npy arrays and a text summary. Each page shows the model name and dataset name at the top so you never lose track of which run produced which output.
| Task | Script | Output |
|---|---|---|
| Classification | visualize_cls.py |
per-sample 3D HTML, confusion matrix, gallery index.html |
| Semantic segmentation | visualize_seg.py |
per-block GT/Pred side-by-side, per-class IoU summary |
| Part segmentation | visualize_partseg.py |
per-shape GT/Pred side-by-side, per-category mIoU summary |
Example — part segmentation with a fine-tuned PointBERT checkpoint:
python scripts/visulization/visualize_partseg.py \
--config cfgs/segmentation/PointBERT/ShapeNetParts/pointbert_partseg.yaml \
--ckpt experiments/ShapeNetParts/pointbert_partseg/ckpt-best-seg.pth \
--num_vis 30 \
--out_dir experiments/ShapeNetParts/pointbert_partseg/vis(The PointBERT GT-vs-Pred image shown at the top of this README was generated by this command.)
Open any .html in your browser for orbit/zoom/pan 3D interaction. The companion .npy files ([N, 5] = x y z gt pred for part seg; [N, 8] with RGB for semseg) are there for downstream analysis or custom plots.
This is the feature that turns LIDARLearn from a training library into a full research pipeline. Run a sweep, call one report script, and walk away with a fully-typeset LaTeX PDF + CSV + Markdown ready to drop into a paper submission. No hand-built tables, no copy-pasting numbers into Overleaf, no formatting regressions when you rerun.
See the actual output: docs/model_comparison_tables.pdf — generated end-to-end by this library from raw training logs.
All report generators live in scripts/reports/ and emit a standalone,
pdflatex-compilable .tex (plus CSV + Markdown) into <exp_dir>/latex/:
| Script | Produces |
|---|---|
classification_comparison_report.py |
CV classification table (one combined table) |
classification_comparison_split_report.py |
CV classification table split by SSL init source |
partseg_comparison_report.py |
ShapeNetParts part-seg comparison table |
semseg_comparison_report.py |
S3DIS semantic-seg summary + optional per-class IoU table (--per_class) |
fewshot_comparison_report.py |
Point-BERT-style few-shot table (5/10-way × 10/20-shot) |
friedman_significance_report.py |
Friedman + Nemenyi + Wilcoxon + CD-diagram report (written to <exp_dir>/friedman/) |
model_citations_report.py |
Standalone citation paragraph for every supported method |
Every generator accepts --use_citation_in_tables (inline \citep{} in each
cell). When omitted (default) a single \paragraph{Methods.} is prepended
above the tables. Bib keys are resolved from scripts/reports/references.bib.
Best-in-column highlighting. Every numeric column is scanned per-report
and the best value (highest for accuracy/IoU/F1/recall/precision, lowest for
param count and epoch time) is automatically wrapped in \textbf{} in the
LaTeX output and **…** in the Markdown output, so the top method stands
out at a glance. Ties are all bolded.
Example reports (fully compiled PDFs + source .tex / .csv / .md
from real benchmark runs) are shipped under experiments/<dataset>/latex/:
| Dataset | Task | Example output |
|---|---|---|
| HELIALS | Classification | experiments/HELIALS/latex/ — model_comparison_tables.{tex,pdf,csv,md} |
| S3DIS | Semantic segmentation | experiments/S3DIS/latex/ — s3dis_comparison.{tex,pdf,csv,md} |
| ShapeNetParts | Part segmentation | experiments/ShapeNetParts/latex/ — shapenetparts_comparison.{tex,pdf,csv,md} |
Each folder includes references.bib (auto-copied from scripts/reports/references.bib) so the .tex files compile standalone with pdflatex + bibtex — drop them straight into a paper submission.
LIDARLearn/
├── cfgs/ # YAML configs (classification, segmentation, fewshot, pretrain)
├── datasets/ # Dataset loaders (ModelNet, ShapeNet, S3DIS, STPCTLS, HELIALS)
│ # + augmentation.py (shared pointcloud transforms)
├── models/ # Backbones, SSL methods, PEFT adapters, seg wrappers
├── tools/ # Training runners (finetune, pretrain, seg, fewshot)
├── utils/ # Config, logging, metrics, checkpoint, parser
├── scripts/ # Training sweeps (train_*.sh), visualisation, smoke tests
│ └── reports/ # LaTeX report generators (classification, partseg, semseg,
│ # fewshot, citations, Friedman significance) + shared references.bib
├── extensions/ # CUDA C++ extensions (pointnet2_ops, chamfer_dist, dela_cutils)
├── tests/ # pytest suite (L1-L4)
├── docs/ # Paper sources and documentation
└── main.py # Unified entry point
pytest tests/ # full suite (~4 min on a single GPU)
pytest tests/test_configs.py # L1: every YAML parses and builds
pytest tests/test_forward_shapes.py # L2: every model forward-passes correctlyHelp wanted — these are the priorities we're actively looking for community contributions on:
- More tests & benchmark feedback, especially on long-running datasets such as S3DIS (semantic segmentation runs take many hours per config). If you have cluster time, please share logs, mIoU numbers, and any hyperparameter adjustments that worked for you.
- Support arbitrary point dimensions (D > 3) — current loaders mostly consume XYZ (plus RGB for a few datasets). Remote-sensing pipelines routinely need per-point intensity, return number, number of returns, classification, scan angle, GPS time, multi-spectral bands, etc. Goal: configurable input channels end-to-end (LAS/LAZ preprocessors → datasets → model
in_channels) without per-model forks. - Reproduce and compare against paper-reported values. For each model + dataset pair, we want a side-by-side table of LIDARLearn's numbers vs. the original paper's. Mismatches are opportunities to fix configs, augmentation pipelines, or training schedules.
- Add new models — recent point cloud backbones, SSL methods, and PEFT strategies that aren't in the matrix above. Follow the three-step contribution workflow below.
- Add 3D point cloud registration support — task heads, dataset loaders (e.g., 3DMatch, KITTI odometry, ModelNet-registration), metrics (rotation/translation error, RRE/RTE), and reference methods (PPFNet is already vendored but not yet wired as a registration task).
- Improve and optimize the code — profile training/inference hot paths, reduce redundant tensor copies, add
torch.compile/ AMP support where it helps, cut memory overhead in large-scene seg runs, and clean up any remaining upstream-vendored code paths. PRs that shave GPU hours are especially welcome. - Windows and macOS support — the project is currently developed and tested on Linux only. Windows needs PowerShell/
.batequivalents forextensions/install_extensions.shand thescripts/train_*.shsweep runners (or a WSL2 setup guide). macOS has no CUDA, so support would mean a CPU/MPS mode that gates the ~6 CUDA-only models (PointTransformerV2/V3, P2P, DeepGCN, DELA, RandLA-Net) ontorch.cuda.is_available()and skips the CUDA extensions build. PRs adding either platform — with install instructions and at least one model verified end-to-end — very welcome. - Full documentation of every model and its hyperparameters — a per-model reference page (markdown under
docs/models/or a single consolidated table) covering: what the model does in 1-2 sentences, its YAMLmodel:block keys, what each hyperparameter controls, recommended defaults per dataset family (ModelNet40 / ShapeNet / STPCTLS / HELIALS / S3DIS), memory/compute footprint, and any model-specific gotchas (e.g., input channel assumptions, required preprocessors, supported task heads). - Any community proposal that improves LIDARLearn — new features, API refinements, documentation, tooling, CI, visualization helpers, or anything else. Open an issue first to discuss scope, then send a PR.
Contributions are welcome! Whether you're adding a new backbone, a new dataset loader, a PEFT strategy, or fixing a bug, the workflow is:
- Fork the repository and create a feature branch (
git checkout -b feat/my-model). - Add your model in three steps — create
models/<name>/<name>.pywith a class inheritingBasePointCloudModel(orBaseSegModel) and decorated with@MODELS.register_module(), register it inmodels/__init__.py, and provide at least one YAML config undercfgs/. - Run the test suite —
pytest tests/must pass. New models are automatically picked up by the L1 config-parse and L2 forward-shape tests. - Follow the style — keep YAML configs
_base_-inherited, document non-obvious hyperparameters, and include the paper citation in your config header. - Open a PR describing the change, the reference paper/repository, and any benchmark numbers if available.
For bugs, feature requests, or questions, please open an issue on GitHub with a minimal reproducer (config + command + error log).
LIDARLearn is released under the MIT License — see LICENSE.
Individual model implementations retain their original licences (MIT or Apache-2.0). Full attribution is provided in THIRD_PARTY_NOTICES.md.
LIDARLearn builds on the outstanding work of the open-source point cloud community. Full per-model attribution — author, venue, original repository, and upstream licence — lives in THIRD_PARTY_NOTICES.md; docs/point_cloud_methods.csv has paper titles and venues.
Canonical BibTeX entries for every cited work (datasets, supervised backbones, SSL methods, PEFT strategies) are kept in scripts/reports/references.bib. Every LaTeX report produced by scripts/reports/*.py resolves its \citep{} commands against this file, and auto-copies references.bib into <exp_dir>/latex/ so the generated .tex compiles end-to-end with:
cd <exp_dir>/latex && pdflatex <report>.tex && bibtex <report> && pdflatex <report>.tex && pdflatex <report>.texFramework inspiration: Pointcept, OpenPoints, Torch-Points3D, Learning3D.
KAN layers used by PointKAN and KAN-DGCNN are built on top of efficient-kan (MIT) — a fast, drop-in replacement for the original Kolmogorov–Arnold Network implementation.
If you use LIDARLearn in your research, please cite:
@misc{ohamouddou2026lidarlearnunifieddeeplearning,
title={LIDARLearn: A Unified Deep Learning Library for 3D Point Cloud Classification, Segmentation, and Self-Supervised Representation Learning},
author={Said Ohamouddou and Hanaa El Afia and Abdellatif El Afia and Raddouane Chiheb},
year={2026},
eprint={2604.10780},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.10780},
}Preprint: https://arxiv.org/abs/2604.10780
Issues and pull requests are welcome on GitHub.
For direct inquiries, reach the maintainer at said.ohamouddou1998@gmail.com.
