SCPP: A Unified Python Library for Soft Clustering

arXiv:2607.19620 · cs.LG, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SCPP: A Unified Python Library for Soft Clustering".

Jane: The paper was written by Kiyan Rezaee, Morteza Ziabakhsh, Artin Bahrampour, Seyed Mohammad Ghoreishi, Asal Khaje et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re cracking open a fresh one from arXiv — it’s called “SCPP: A Unified Python Library for Soft Clustering.” Jane, I’ll be honest, the title alone got me curious. Soft clustering — that’s not something we hear every day.

Jane: It’s a great place to start, Tom. So, hard clustering is what most of us know — think of k-means, where every data point gets forced into exactly one group, like assigning each student to one homeroom. Soft clustering says, hey, why not let a student spend part of their day in multiple homerooms? Each point gets a fraction of membership in every cluster.

Tom: Right, so instead of a hard label, you get a vector of probabilities. That’s huge for things like overlapping communities in social networks, or documents that touch on several topics at once.

Jane: Exactly. And the paper’s point is that the Python ecosystem for this is a mess. There are libraries for fuzzy clustering, separate ones for probabilistic models, others for graph-based methods — and they all speak different APIs. If you want to compare a fuzzy c-means model to a graph neural network approach, you’re basically writing custom glue code for each one.

Tom: That sounds painful. So what does SCPP actually do about it?

Jane: It unifies forty different soft clustering algorithms under one scikit-learn-style interface. You call fit, you call predict, you get a membership matrix out. Whether it’s a classic fuzzy method or a deep learning model, the workflow looks identical.

Tom: Forty algorithms? That’s a serious catalog. And they span everything from fuzzy c-means to graph neural networks?

Jane: You’ve got it. Foundational stuff like FCM and Gaussian mixture models, then modern variants, graph-based community detection like BigCLAM, even deep methods like NOCD. All under one roof.

Tom: So for a researcher, this could be a game-changer. Instead of wrestling with five different libraries, you just pick your algorithm from a menu.

Jane: And that’s exactly the pitch. The authors want to make soft clustering as easy to use and benchmark as hard clustering already is with scikit-learn.

Tom: I love that ambition. But I’m already wondering — how do you actually evaluate these things fairly when they’re so different under the hood? That feels like a big piece of the puzzle.

Jane: It is, and the paper has a whole framework for that. But let’s save that for the next segment — we’ve got a lot more ground to cover.

Summary: Tom: We’re back with “SCPP: A Unified Python Library for Soft Clustering,” and Jane, you teased that the evaluation story is a big deal. Let’s dig into what the paper actually delivers.

Jane: So beyond the forty algorithms, the real meat is the benchmarking layer. The authors built twenty benchmark datasets, twelve clustering quality metrics, and dedicated tools for measuring runtime, memory, and scalability. All of it plugs into the same unified interface.

Tom: Twenty datasets — that’s a solid test bed. And twelve metrics? That’s a lot of ways to judge a clustering result.

Jane: It is, and they split them into two buckets. There are the hard metrics you already know — silhouette score, adjusted Rand index, normalized mutual information. Those work on the crisp labels you get by taking the argmax of the membership matrix. But then there are soft-specific metrics that operate directly on the membership values themselves.

Tom: Like what?

Jane: Partition coefficient, partition entropy, Xie-Beni index. These measure how fuzzy or how crisp the partition is, how much uncertainty is baked into the assignments. A hard clustering algorithm can’t even compute those — you need the full membership matrix.

Tom: That’s a really important distinction. So the paper isn’t just saying “here are more algorithms,” it’s saying “here’s a fair way to compare them all.”

Jane: Precisely. And they show it with a benchmark table — eight algorithms on six synthetic datasets, reporting external metrics, internal hard metrics, soft partition metrics, and computation time. You can see, for example, that EntropyFCM gets the best adjusted Rand index on average, while GMM has the best partition coefficient.

Tom: And the runtime numbers are interesting too — FCM is the fastest at under four milliseconds, while GK takes over thirty.

Jane: Right, but here’s the thing — the paper isn’t declaring a single winner. The point is that now you can actually see these trade-offs side by side without writing custom evaluation code for every method.

Tom: That’s the kind of infrastructure the field needs. I mean, how many papers compare fuzzy clustering to graph-based methods and just eyeball it?

Jane: Exactly. And they even have a scalability module that shows how runtime and memory grow as you increase the dataset size. For FCM, memory grows linearly — about zero point two six kilobytes per sample, with an R-squared above zero point nine nine nine.

Tom: That’s a clean result. So the benchmarking story is solid. But what about the practical side — is this actually usable in a real project, or is it just a research toy?

Jane: That’s the question, isn’t it? And the paper has a whole case study about integrating SCPP into a real migration tool called Mo2oM. Let’s talk about that in the next segment.

Improvements: Tom: We’re still on “SCPP: A Unified Python Library for Soft Clustering,” and Jane just mentioned the Mo2oM case study. That sounds like a real-world test, which I love.

Jane: Mo2oM is a tool that helps break down monolithic software into microservices. Originally, it was hardwired to one specific clustering method — NOCD, a graph neural network approach. If you wanted to try a different soft clustering algorithm, you had to rewrite the integration code.

Tom: That’s the classic problem — one library, one method, no flexibility.

Jane: Exactly. So the authors integrated SCPP into Mo2oM, and suddenly they could swap between NOCD, Bayesian non-negative matrix factorization, and BigCLAM just by changing a parameter. The core migration logic stays untouched.

Tom: And the results? Did the different methods actually change the outcome?

Jane: They did, and the paper shows it. On the JPetStore dataset, BayesianNMF gets a perfect structural similarity score of one point zero, while NOCD gets zero point eight zero six. The interface coupling penalty is zero point zero for BayesianNMF but zero point eight zero six for NOCD. These are very different migration recommendations.

Tom: So the choice of clustering algorithm genuinely matters for the quality of the microservice decomposition.

Jane: That’s the whole point. And before SCPP, you couldn’t even run that comparison without significant engineering effort. Now it’s a few lines of code.

Tom: That’s a compelling demonstration. But I’m curious about the engineering side — what does it take to add a new algorithm to this framework?

Jane: The paper says you just implement the standard interface — fit, predict, predict proba — and you immediately get access to all the benchmarking, evaluation, and visualization tools. The dependencies are lightweight too, mostly NumPy and SciPy, with PyTorch only needed for graph and deep learning methods.

Tom: That’s smart. Keep the core light, make the heavy stuff optional.

Jane: And they’ve got the software engineering chops to back it up — two hundred forty-one unit tests across forty-one modules, continuous integration, and a public repository on GitHub.

Tom: That’s real research-grade infrastructure. Not just a script someone threw together.

Jane: Exactly. And they’ve also thought about the broader Python ecosystem — the appendix has detailed guides for preprocessing tabular data, text, images, and graphs before feeding them into SCPP. So it’s not just a clustering library; it’s a full workflow tool.

Tom: I love that. It makes the library approachable for people who aren’t fuzzy clustering experts.

Jane: That’s the goal. And the implications go beyond just convenience — when you make it easy to compare methods fairly, you make it easier to discover which algorithms actually work best for which problems.

Tom: So what’s the bigger picture here? Where does this leave the field?

Jane: That’s what we’ll wrap up with in the final segment.

Conclusion: Tom: Alright, we’re wrapping up our look at “SCPP: A Unified Python Library for Soft Clustering.” Jane, give us the final take.

Jane: The big picture is that this library removes a real barrier. Soft clustering has been fragmented across incompatible packages, and SCPP brings forty algorithms under one consistent, scikit-learn-compatible interface. That means researchers can spend their time on the science, not on glue code.

Tom: And the benchmarking layer is the unsung hero. Twenty datasets, twelve metrics, runtime and memory profiling — all standardized. That’s what makes fair comparisons possible.

Jane: The Mo2oM case study really drove it home for me. Different clustering methods gave genuinely different microservice decompositions, and SCPP made it trivial to see that. That’s the kind of practical impact that matters.

Tom: And with two hundred forty-one tests and a clean API, it’s built to last. This isn’t a weekend project — it’s a serious contribution to the scientific Python ecosystem.

Jane: For sure. And the fact that it’s open source under the MIT license means anyone can extend it. Add a new algorithm, and it instantly works with all the existing evaluation tools.

Tom: So who benefits most from this?

Jane: Anyone working with overlapping clusters — bioinformatics, natural language processing, social network analysis, software engineering. If your data has ambiguity, soft clustering is the right tool, and now it’s actually easy to use.

Tom: And that’s the kind of progress we love to see. A tool that makes good science easier to do.

Jane: Absolutely. We’ll be watching to see what algorithms get added next and how the community adopts it.

Tom: Thanks for joining us, everyone. Next time, we’ll be looking at a fresh paper on temporal graph benchmarks — see you then.

Kiyan Rezaee, Morteza Ziabakhsh, Artin Bahrampour, Seyed Mohammad Ghoreishi, Asal Khaje, Ali Sajedifar, Manny Chalak, Ava Zerafatangiz, Sadegh Eskandari

cs.LG, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: 4 pages

Code: https://github.com/soft-clustering/soft-clustering

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

Key concepts

Soft Clustering
Unlike hard clustering where every data point is forced into one group, soft clustering allows points to have a fractional membership in multiple clusters. This results in a vector of probabilities and is useful for modeling overlapping communities or documents that touch on several topics.
Hard Clustering
This is the traditional method where each data point is assigned to exactly one cluster, similar to assigning a student to a single homeroom. It provides a crisp, single label for every input data point.
SCPP Library
A unified Python library that integrates forty different soft clustering algorithms. It provides a consistent, scikit-learn-style workflow—call fit and predict—allowing researchers to compare complex methods easily without writing custom glue code.
Benchmarking
The paper provides a standardized framework using twenty benchmark datasets and twelve metrics. This allows for fair comparison between soft clustering algorithms, including hard metrics like the adjusted Rand index and soft-specific measures like partition entropy.

Terminology

Summary

Summary

SCPP (Soft Clustering Python Package) is an open-source Python framework for soft clustering, which assigns each observation a vector of fractional memberships rather than a single cluster label. The paper states: "SCPP establishes a canonical, scikit-learn-compatible estimator interface that standardizes model training, prediction, membership representation, evaluation, and benchmarking across heterogeneous soft clustering methods, including fuzzy, probabilistic, graph-based, matrix factorization, and deep learning methods."

The framework currently integrates 40 representative algorithms and provides a comprehensive benchmarking comprising datasets, clustering quality metrics, and standardized runtime, memory, and scalability evaluation. The source code is publicly available at https://github.com/soft-clustering/soft-clustering.

The paper identifies a problem: "Despite extensive methodological advances, the software ecosystem for soft clustering remains fragmented. Existing implementations are distributed across independent Python libraries that differ in algorithm coverage, APIs, parameter conventions, membership representations, and benchmarking support." Most libraries focus on a single family of methods rather than providing a common abstraction across heterogeneous paradigms, hindering reproducibility, systematic benchmarking, and fair empirical comparison.

The contributions of the work are fourfold:

  1. A unified software abstraction for soft clustering: SCPP introduces a scikit-learn-compatible estimator interface that enables diverse soft clustering algorithms to be trained, evaluated, and compared through the same workflow, independent of their underlying methodology.

  2. Comprehensive algorithmic coverage within a common framework: The library unifies 40 representative soft clustering algorithms spanning the principal methodological families under a consistent API.

  3. An integrated benchmarking and reproducibility ecosystem: SCPP provides standardized datasets, clustering quality metrics, runtime, memory, and scalability benchmarking modules that facilitate fair and reproducible empirical evaluation.

  4. Research-grade software engineering infrastructure: "To support reliable scientific software, the package is distributed through PyPI under the MIT license and is accompanied by comprehensive documentation, extensive examples, automated testing comprising 241 unit tests across 41 modules, and continuous integration workflows."

Design and Implementation: The central design principle is that heterogeneous soft clustering methods should expose a common software abstraction despite substantial differences in their mathematical formulations and optimization procedures. This abstraction standardizes the complete lifecycle of a soft clustering estimator, including model configuration, training, prediction, membership representation, evaluation, and benchmarking.

The architecture is organized into five software layers: end users interact exclusively through a canonical, scikit-learn-compatible estimator interface; the underlying implementation encapsulates multiple families of soft clustering algorithms within a common API; a dedicated benchmarking layer provides standardized evaluation infrastructure; and the framework is complemented by documentation, automated testing, and seamless integration with the scientific Python ecosystem.

Every estimator implements the standard methods fit, predict, predict proba, and fit predict, while exposing consistent outputs through attributes such as membership, labels, and, when available, cluster centers and weights. The paper provides an example workflow:


from sklearn.datasets import load iris

from soft clustering import FCM

X, = load iris(return X y=True)

model = FCM(n clusters=3, random state=0)

model.fit(X)

U = model.membership

labels = model.labels

Benchmarking: The benchmarking layer comprises 20 benchmark datasets, 12 clustering quality metrics, standardized experiment runners, and dedicated modules for runtime, memory, and scalability analysis. Because every estimator conforms to the same software abstraction, identical evaluation pipelines can be executed across all supported methods.

Algorithm coverage: The 40 algorithms span the following categories:

  • Foundational Methods: FCM, GMM, PCM, Gustafson–Kessel, Gath–Geva

  • Classical Extensions: Rough K-Means, Latent Dirichlet Allocation (LDA), Probabilistic Latent Semantic Analysis (PLSA), Subtractive Clustering

  • Modern Variants: Kernelized FCM, Fuzzy-Possibilistic C-Means, Adaptive FCM with Spatial Regularization, Soft Kernel Spectral Clustering, Soft DBSCAN-GM, Fuzzy Competitive Clustering, Multivariate Beta Mixture Models, Beta-Gaussian Mixture Models, Evidential C-Means, EVCLUS

  • Graph and Community Detection: BigCLAM, Bayesian Nonnegative Matrix Factorization, Mixed Membership Stochastic Block Models, Gumbel-Softmax Community Detection, NOCD, Deep Modularity Networks (DMoN)

  • Ensemble and Advanced Methods: Soft Clustering Ensembles, Adaptive FCM with Graph Embedding, Robust Projected Fuzzy K-Means, Deep Fuzzy K-Means, Federated Fuzzy Clustering, Semi-supervised Fuzzy Clustering, Fuzzy Subspace Clustering, Fuzzy Soft Set Clustering

  • Specialized Techniques: Entropy C-Means, CAF-HFCM, Collaborative Annealing FCM, Fuzzy Color Clustering, Word-Based Soft Clustering, Similarity-Based Soft Clustering, KMART

Datasets: The benchmark dataset collection includes 20 datasets organized into three groups: real-world (iris, wine, digits, breast cancer, olivetti faces), synthetic (blobs, moons, circles, anisotropic blobs, varied blobs, high dimensional blobs), and OpenML (glass, vehicle, ecoli, yeast, segment, satimage, letter, pendigits, optdigits).

Evaluation Framework: SCPP provides a unified evaluation framework for both hard and soft clustering methods. In addition to conventional clustering metrics, the framework includes soft-clustering-specific measures that directly operate on the membership matrix U. These include partition coefficient, modified partition coefficient, partition entropy, fuzzy hypervolume, Xie-Beni index, fuzzy compactness, and fuzzy separation. For hard cluster assignments, SCPP integrates widely used internal and external evaluation metrics from scikit-learn, including silhouette, Calinski-Harabasz, Davies-Bouldin, adjusted Rand index (ARI), and normalized mutual information (NMI).

Computational Performance Benchmarking: The framework evaluates runtime, scalability, and memory consumption using a unified interface. The runtime benchmark repeatedly executes the training procedure and reports average execution time and standard deviation across multiple runs. Scalability is assessed on progressively larger subsets of the same dataset, with sample sizes ranging from 100 to 10,000 observations. Memory profiling measures resident memory usage before and after model fitting, computes memory overhead, and records peak observed memory footprint.

Testing Infrastructure: The testing framework consists of 241 unit and integration tests distributed across 41 dedicated test modules, validating every one of the 40 clustering algorithms. The verification strategy includes output validation, mathematical consistency (non-negativity of memberships, normalization constraints, parameter ranges), model state verification, parameter robustness, edge cases and failure modes, and reproducibility (repeated executions with identical random seeds produce identical outputs).

Mo2oM Migration Case Study: The paper demonstrates a real-world use case integrating SCPP into Mo2oM, a monolithic-to-microservice migration pipeline. Originally hardwired to a single graph community detection method (NOCD), after integration Mo2oM can select from multiple graph-based soft clustering methods (NOCD, BayesianNMF, BIGCLAM) through a unified API, enabling pluggable clustering method selection, rapid experimentation, and preserving soft membership semantics for overlap-aware service decomposition.

Integration with Python Ecosystem: The paper provides guidance for integrating SCPP with the PyData stack, covering tabular data preprocessing (pandas, scikit-learn), text data embedding (sentence-transformers, TF-IDF), image feature extraction (torchvision ResNet), graph/relational data construction (networkx, scipy sparse matrices), and evaluation metrics. The recommended end-to-end workflow is: (1) load and preprocess raw data into a feature matrix X or affinity matrix A; (2) fit the model with the unified SCPP interface; (3) extract and interpret the membership matrix U to obtain hard labels, confidence scores, entropy, and cluster masses; (4) evaluate comprehensively using both soft-specific internal metrics (PC, PE, XB) and standard external metrics (ARI, NMI).

The framework adopts a lightweight dependency strategy, relying primarily on NumPy and SciPy while requiring PyTorch only for graph-based and deep learning models. The architecture is designed to facilitate long-term extensibility and integration with the broader scientific Python ecosystem.

Improvements for AI systems

Based on the SCPP paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

  • Implement a scikit-learn-compatible estimator interface with standardized fit, predict, predict proba, and fit predict methods across all clustering algorithms

  • Standardize membership matrix output (U ∈ [0,1](n×K)) with consistent attribute naming (membership, labels, centers)

  • Add automatic parameter validation for fuzzification coefficients, convergence tolerances, and initialization strategies

  • Integrate 40 algorithms across 7 methodological families: fuzzy (FCM, PCM, GK), probabilistic (GMM, LDA, PLSA), graph-based (BigCLAM, NOCD, DMoN), matrix factorization (BayesianNMF), deep learning, ensemble, and specialized methods

  • Implement a pluggable architecture allowing new algorithms to inherit the unified interface and immediately work with existing benchmarking infrastructure

  • Add soft-clustering-specific metrics that operate directly on membership matrices: partition coefficient, partition entropy, fuzzy hypervolume, Xie-Beni index, fuzzy compactness, fuzzy separation

  • Integrate standard hard clustering metrics (ARI, NMI, silhouette, Calinski-Harabasz, Davies-Bouldin) for defuzzified labels

  • Provide a unified evaluation interface that automatically computes all applicable metrics from labels, memberships, and prototypes

  • Implement standardized runtime benchmarking with repeated fits, mean/std reporting, and prediction time measurement

  • Add memory profiling with tracemalloc for peak memory, memory overhead, and resident memory tracking

  • Create scalability analysis across sample sizes (100 to 10,000) measuring both runtime and memory growth

  • Implement 241 unit tests across 41 modules validating: output dimensions, mathematical invariants (membership normalization, non-negativity), parameter robustness, edge cases, and deterministic behavior with fixed random seeds

  • Add continuous integration workflows and comprehensive documentation for reproducible experimentation

  • Build preprocessing pipelines for tabular data (scaling, encoding, imputation), text data (LLM embeddings, TF-IDF), image data (ResNet features), and graph data (affinity matrix construction)

  • Integrate with PyData ecosystem: pandas/polars, scikit-learn, sentence-transformers, torchvision, networkx, faiss

  • Compare 40+ clustering algorithms with identical code — swap between FCM, GMM, NOCD, and BayesianNMF by changing one line

  • Run complete benchmark suites (quality, runtime, memory, scalability) with a single API call

  • Evaluate soft clustering quality using both fuzzy-specific metrics (PC, PE, XB) and standard metrics (ARI, NMI)

  • Reproduce experiments exactly with deterministic random seeds and standardized evaluation protocols

  • Handle heterogeneous data types (tabular, text, image, graph) through unified preprocessing pipelines

  • Extract richer insights from membership matrices: confidence scores, uncertainty entropy, cluster mass, and ambiguous sample identification

  • Deploy in production with validated numerical stability, edge-case handling, and predictable failure modes

  • Extend with new algorithms by implementing a single estimator interface — new methods immediately work with benchmarking, evaluation, and visualization tools

  • Integrate into existing pipelines (e.g., Mo2oM migration tool) with pluggable clustering method selection

  • Run automated testing with 241 unit tests ensuring mathematical correctness and robustness

  • Improve software modularization by identifying overlapping components with cross-cutting responsibilities (as demonstrated in the Mo2oM case study)

  • Enhance community detection in social networks with soft memberships capturing overlapping community structures

  • Improve topic modeling with semantic-aware clustering that preserves uncertainty in document assignments

  • Process datasets from 100 to 10,000+ samples with linear memory scaling (0.26 KB per sample for FCM)

  • Handle sparse matrices and disconnected graph structures without failure

  • Validate mathematical constraints automatically: membership normalization (sum to 1), non-negativity, parameter ranges

  • Provide confidence-calibrated outputs where high-entropy points are flagged for manual review or active learning

Abstract

In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a canonical, scikit-learn-compatible estimator interface that standardizes model training, prediction, membership representation, evaluation, and benchmarking across heterogeneous soft clustering methods, including fuzzy, probabilistic, graph-based, matrix factorization, and deep learning methods. The framework currently integrates 40 representative algorithms together with a comprehensive benchmarking comprising datasets, clustering quality metrics, and standardized runtime, memory, and scalability evaluation. SCPP further provides extensive documentation, practical examples, automated testing, and seamless integration with the scientific Python ecosystem, enabling reproducible experimentation and straightforward extension with new algorithms. The source code is publicly available at https://github.com/soft-clustering/soft-clustering.

Sources

Related papers