--- license: cc-by-nc-4.0 library_name: torch-pointcloud tags: - point-cloud - 3d - pytorch - torch-pointcloud - concerto - segmentation datasets: - scannet base_model: torch-pointcloud/concerto-large.pretrain.pointcept model-index: - name: concerto-large-lp.scannet20.pointcept results: - task: type: point-cloud-segmentation dataset: name: ScanNet (20 classes) type: scannet metrics: - name: mIoU type: mean_iou value: 78.59 --- # Model card for concerto-large-lp.scannet20.pointcept A Concerto point cloud segmentation model (joint 2D-3D representation encoder). Trained on ScanNet (20 classes). > **Non-commercial.** These weights are released by [Pointcept/Concerto](https://github.com/Pointcept/Concerto) under CC BY-NC 4.0 and may be used for research and evaluation only. ## Model Details - **Model Type:** Point cloud semantic segmentation - **Model Stats:** - Params (M): 207.7 - Input channels: 9 - Classes: 20 - Features: 1728 - **Dataset:** ScanNet (20 classes) - **Metrics:** mIoU 78.59 (reference 77.5) - **Paper:** [Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations](https://arxiv.org/abs/2510.23607) - **Converted from:** [Pointcept/Concerto](https://github.com/Pointcept/Concerto) (CC-BY-NC-4.0) - **Library:** [torch-pointcloud](https://github.com/arthurdjn/pytorch-pointcloud) ## Install ```bash pip install torch-pointcloud ``` This checkpoint also needs `spconv` and `flash-attn`, which need a build matching your torch and CUDA: see the [installation guide](https://pytorch-pointcloud.org/installation/). ## Usage ```python import torch import torch_pointcloud as tp from torch_pointcloud.utils.data import collate model, info = tp.create_model( "concerto-large-lp.scannet20.pointcept", task="segmentation", pretrained=True, return_info=True, ) model = model.cuda().eval() # GPU-only kernels # synthetic sample with the keys a dataset provides num_points = 8192 sample = { "pos": torch.randn(num_points, 3), "color": torch.rand(num_points, 3) * 255, "normal": torch.randn(num_points, 3), "segment": torch.zeros(num_points, dtype=torch.long), "instance": torch.zeros(num_points, dtype=torch.long), } data = info["transform"](sample) data = collate([data]) data = {key: value.cuda() for key, value in data.items()} with torch.no_grad(): logits = model(data.get("x"), data["pos_grid"], data["batch"]) ``` ## Feature extraction ```python with torch.no_grad(): features = model.forward_features(data.get("x"), data["pos_grid"], data["batch"]) model.reset_classifier(num_classes=0) with torch.no_grad(): features = model(data.get("x"), data["pos_grid"], data["batch"]) # (N, 1728) ``` ## Citation ```bibtex @article{concerto2025, title = {Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations}, author = {Yujia Zhang and Xiaoyang Wu and Yixing Lao and Chengyao Wang and Zhuotao Tian and Naiyan Wang and Hengshuang Zhao}, journal = {arXiv preprint arXiv:2510.23607}, year = {2025} } @inproceedings{dai2017scannet, title = {ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes}, author = {Angela Dai and Angel X. Chang and Manolis Savva and Maciej Halber and Thomas Funkhouser and Matthias Nießner}, booktitle = {CVPR}, year = {2017} } @software{dujardin2026pytorchpointcloud, author = {Arthur Dujardin}, title = {PyTorch PointCloud}, year = {2026}, doi = {10.5281/zenodo.22159632}, url = {https://github.com/arthurdjn/pytorch-pointcloud}, } ```