TabSOM: A tabular-to-image encoding method based on self-organizing maps

arXiv:2608.13513 · cs.CV, cs.LG · Submitted 2026-08-13 · Read on arXiv

David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara, Francisco J. Lara-Abelenda, Luis Zhinin-Vera, Diego H. Peluffo-Ordóñez

Rey Juan Carlos University · Universidad Francisco de Vitoria · University of Castilla-La Mancha · Yachay Tech University

cs.CV, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 65/100

The gist: TabSOM: A tabular-to-image encoding method based on self-organizing maps Abstract Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of

Terminology

Summary

TabSOM: A tabular-to-image encoding method based on self-organizing maps

Abstract

Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction method (e.g., t-SNE, UMAP, PCA). However, they encode only the marginal value of each feature and discard information about feature relationships. We propose TabSOM, a tabular-to-image encoding built on the Self-Organizing Map (SOM), which provides: (i) a spatial layout in which every input feature occupies a fixed canvas position derived from its component plane via collision-free Hungarian assignment; and (ii) a graph that captures pairwise feature relationships derived from the SOM component planes. The resulting image stacks two multi-scale node channels: one encodes feature values at fixed scales, while the other encodes pairwise feature interactions as spatial connections between related features. Two SOM-derived interpretability approaches are introduced: a prototype-inspired partial dependence plot and a class-separation importance score. Benchmarked against twelve existing tabular-to-image methods across public binary-classification datasets, TabSOM ranks first or second on every dataset and achieves the lowest variance of any method evaluated. Interpretability obtained with TabSOM was validated against Random Forest, XGBoost, and SHAP, the class-separation score shows reasonable agreement with established baselines on the top-ranked features while capturing complementary structural information from input data. These results demonstrate that TabSOM provides an effective and interpretable approach for applying deep learning architectures to tabular data, bridging the performance–interpretability gap in this domain.

Introduction

Tabular data are among the most widely used data formats for representing structured information. It is characterized by a table-like format, with rows corresponding to samples and columns to features. While Deep Learning (DL) models, such as Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have achieved remarkable results on data characterized by high spatial or temporal correlations (e.g., images and audio), tabular data generally exhibit weak interdependencies among features. As a result, DL architectures often show reduced effectiveness on predictive tasks involving tabular data. This has fostered recent research into encoding methods, which convert tabular data into images to leverage the predictive performance of CNNs and pretrained ViTs.

Several tabular-to-image methods have been proposed in the literature, ranging from parametric to non-parametric approaches. Most of these methods, such as DeepInsight, TINTO, REFINED, project features into two-dimensional through dimensionality reduction methods. They generate images in which spatial proximity reflects feature similarity, with neighboring pixels representing features with some similarity. Tabular-to-image methods have demonstrated competitive performance in several domains such as healthcare, indoor localization, and the internet of things among others.

The performance of tabular-to-image encoding methods directly depends on how features are spatially arranged and how their values are encoded. Existing methods have addressed this mainly through dimensionality reduction methods, including Principal Component Analysis (PCA), kernel PCA, t-distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP). These project the feature space into two-dimensional space, assigning each feature to its nearest pixel on the image canvas. Although they produce effective spatial layouts, the embedding is used solely to determine feature locations. As a result, each feature occupies a fixed position derived from the feature distribution, without explicitly encoding relationships among feature values.

The Self-Organizing Map (SOM), an artificial neural network based on unsupervised and competitive learning, performs a topology-preserving mapping from the input space to a two-dimensional grid of nodes. The topology-preserving property enables that similar inputs are mapped to neighbor nodes on the grid. Then, local relationships in the input space are reflected in spatial proximity on the grid. As a result, SOM has been extensively used for visualization, clustering, and pattern recognition in different domains such as signal processing, healthcare among others. Although SOM has shown excellent results in unsupervised tasks, its application within tabular-to-image methods has been partially explored. Unlike other dimensionality reduction techniques, the SOM preserves an interpretable structure. Each node is associated with a prototype, the grid of prototypes provides a density-weighted summary of the joint feature distribution, and the component planes capture how each feature varies across the map. Although tabular-to-image encoding methods have shown promising predictive performance, their interpretability has received limited attention. The graph structure of the SOM captures the topology of the data while providing interpretability.

In this paper, we propose a tabular-to-image encoding method named TabSOM, which includes pairwise feature interactions as part of the image. We train a SOM on the feature space and extract, for each feature, its component plane, i.e., the spatial distribution of that feature’s values across the trained SOM grid. We derive a canvas position (named anchor) for each feature from its component plane, and then optimize anchor collisions via a Hungarian assignment to grid cells. The SOM’s prototype grid serves as a density-weighted, topologically organized sample of the joint feature distribution, and the CNN’s predictions on those prototypes define a spatial prediction map over the grid. One of the main advantages of TabSOM over other tabular-to-image methods is that the SOM is not used solely to fit the layout but also provides two interpretability methods derived from the grid: (i) global feature ranking; and (ii) prototype-inspired partial dependence curves.

The main contributions of this paper are the following:

  • A novel tabular-to-image encoding method in which feature placement is derived from SOM component planes via anchor estimation and collision-free (Hungarian) assignment.

  • A relational channel that renders pairwise feature interactions derived from the SOM prototype geometry that determines placement.

  • A multi-scale rendering scheme in which the same feature layout is rendered at multiple Gaussian bandwidths as aligned channels, providing a CNN with both precise localization and regional coherence without requiring a single bandwidth choice.

  • An SOM-based interpretability framework that leverages the learned grid structure to provide two complementary explanations: (i) global feature ranking; and (ii) prototype-inspired partial dependence curves.

  • Feature maps that allow the encoding’s class-separability and per-feature spatial assignment to be inspected independently of any downstream classifier.

Related Work

A range of tabular-to-image methods have been proposed in the literature, which are categorized into two main categories: non-parametric and parametric methods.

Non-parametric methods use heuristic algorithms to map features into a predefined image-grid representation without explicitly optimizing the spatial arrangement of features. Representative approaches include the following. BarGraph encodes each tabular sample as a vertical bar chart, where each feature is assigned to a fixed column and the bar height corresponds to the normalized feature value. Binary Image Encoding (BIE) and correlated BIE convert each numerical value to binary strings, then stack these bit-rows to form a two-dimensional matrix where 0 and 1 denote dark and light pixels, respectively. SuperTML encodes feature values as text on an image, using CNNs as character/word-level feature extractors. Tab2Visual encodes each feature as a vertical bar whose width is proportional to the normalized feature value, arranged in rows and columns of bars.

Parametric methods perform the spatial arrangement of features through an optimization-based preprocessing stage. Linear and nonlinear dimensionality reduction techniques (such as PCA, t-SNE, and UMAP) project features from the input space into a low-dimensional coordinate space, with the resulting coordinates defining the feature layout from which the image is built. Among the most popular methods, we find DeepInsight and REFINED, which use t-SNE, kernel PCA and Multidimensional Scaling to obtain a 2D position for each feature, and then map feature values to pixel intensities. TINTO uses PCA and t-SNE to determine feature locations and the resulting coordinates are then transposed, scaled, and rounded to integer values. It incorporates a blurring mechanism to produce smoother image representations, thereby enhancing compatibility with convolutional filters. Fotomics applies the Fourier Transform independently to each feature column and represents the resulting complex coefficients through their real and imaginary components as coordinates in a two-dimensional Cartesian plane. FC-Viz aggregates highly correlated features into clusters and evaluates the relationships between representative features from each cluster. Ant colony optimization determines the spatial organization of the features, while dimensionality reduction methods are employed to compute pixel intensities. Other methods in this category are IGTD and LM-IGTD. They perform a feature-to-pixel assignment as a distance-preservation optimization problem, minimizing the difference feature similarities and distances between samples. LM-IGTD extends IGTD to handle low-dimensional and mixed-type data through an unsupervised stochastic noise generation.

Methods

TabSOM consists of five stages: (1) SOM training and component-plane extraction, (2) feature anchor estimation, (3) collision-free spatial placement, (4) feature relationship graph construction, and (5) multi-channel image rendering.

Self-organizing map training: Let X ∈ RN×F be a tabular dataset with N samples and F features, and let y ∈ 0, 1 N denote the corresponding vector labels. Each feature is first scaled using min-max or z-score normalization. A rectangular SOM with H × W nodes is trained on the normalized data using the standard online SOM update rule. Each node (r, c) has associated a prototype vector wr,c ∈ RF, initialized by sampling samples from the training data. At each training step t, a sample x is drawn at random, and its best matching unit (BMU) is identified as b = arg min∥wr,c − x∥2, and every node’s prototype is updated according to the rule: wr,c ← wr,c + η(t) hσ(t) (r, c, b) (x − wr,c), where the learning rate η(t) and neighborhood radius σ(t) decay exponentially from initial values η0, σ0 to final values η1, σ1 over the course of training, and hσ (r, c, b) = exp(−((r−rb)2+(c−cb)2)/(2σ2)) is a Gaussian neighborhood function over grid coordinates. The grid size H × W is chosen such that the number of cells K = H · W is at least F, ensuring sufficient placement capacity for features.

Component planes and feature anchors: After training, the j-th component plane Φj ∈ RH×W is defined as the j-th coordinate of every node’s prototype vector, Φj (r, c) = [wr,c]j, a smooth spatial field over the grid whose value at (r, c) indicates how high feature j tends to be among samples mapped near that node. Because the SOM update rule pulls neighboring nodes toward similar prototypes, Φj varies gradually across the grid rather than at random. Nodes that are close together on the map represent similar combinations of feature values, then Φj exhibits a coherent region of high intensity rather than several disconnected peaks. This topology preservation is what allows a single scalar feature to be associated with a location on the grid in the first place, rather than only with a value at each node independently. Each feature j is first assigned an anchor position (aj ∈ [0, 1]2), representing its preferred location on the image canvas before collision resolution. Two methods are used: Centroid anchor (the intensity-weighted center of mass of the optionally sharpened component plane) and Mode anchor (the location of the plane’s maximum, optionally refined to sub-pixel precision via parabolic interpolation). In practice, the centroid anchor is preferred when component planes are broad and overlapping (typical of low-to-moderate F), while the mode anchor is preferred when F is large and many planes compete for similar grid regions.

Collision-free feature placement: Anchors computed may coincide or be placed closely, which would cause multiple features to overlap on the rendered image canvas. To solve this, we apply a discrete optimal-assignment step. Let g1,..., gK denote the centers of the K = H × W canvas cells (in normalized [0, 1]2 coordinates). The cost of placing feature j in cell c is Cjc = α ∥gc − aj∥ − β Φ̃j (gc), where the first term penalizes displacement from the feature’s anchor and the second (optional, β ≥ 0) rewards placing the feature where its own component plane is locally strong. The final placement σ: 1,..., F → 1,..., K is obtained by solving the linear assignment problem σ∗ = arg minσ ΣFj=1 Cj,σ(j) via the Hungarian algorithm, guaranteeing that no two features share a cell (zero overlap) while minimizing total displacement from the SOM-derived anchors.

Feature relationship graph: To capture pairwise feature interactions, we construct a feature relationship graph directly from the component planes. Each component plane Φj is flattened to a vector in RHW, and the relationship weight between features i and j is computed as either their Pearson correlation or cosine similarity across the flattened planes: Wij = corr(vec(Φi), vec(Φj)) or Wij = vec(Φi)⊤ vec(Φj) / (∥vec(Φi)∥ ∥vec(Φj)∥). Negative weights are clipped to zero (retaining only positive co-variation), and the diagonal is set to zero. The edge set E is obtained by thresholding: E = (i, j): Wij ≥ τ for a threshold τ (default 0.5–0.6). Because W is derived from the same prototype geometry that determines feature placement, features connected by an edge are typically placed near each other on the canvas.

Multi-channel image rendering: Given the placement pj Fj=1 ⊂ [0, 1]2 (canvas positions) and the edge set E with weights Wij (i,j)∈E, a tabular sample x ∈ RF is rendered into an image I ∈ RHimg×Wimg×C with C channels. We selected C = 3, with two first node channels and a relational edge channel. For each of S Gaussian bandwidths σ1 < σ2 < · · · < σS (default S = 3, e.g. σ ∈ 0.05, 0.10, 0.18, sharp to smooth), one channel is rendered as Is (u, v) = ΣFj=1 xj · exp(−∥(u, v) − pj∥2 / (2σs2)), s = 1,..., S, where (u, v) ranges over pixel centers in [0, 1]2. Each channel is the same spatial layout rendered at a different bandwidth: the sharp channel localizes each feature precisely (supporting fine-grained discrimination), while the smooth channel provides regional coherence over a neighborhood the size of a convolutional kernel’s receptive field (supporting the local-pattern assumptions of CNNs). Because all S channels share the same feature positions, they are spatially aligned, allowing a CNN to jointly exploit multiple scales — analogous to a feature pyramid. We use S = 2 (σ ∈ 0.05, 0.08), reserving the third channel for the relational edge channel. For each edge (i, j) ∈ E, define the segment field as the Gaussian distance from (u, v) to the line segment connecting pi and pj: segij (u, v) = exp(−d((u, v), pi pj)2 / (2σe2)), where d(·, ·) is the perpendicular distance to the segment (clamped to the segment’s endpoints) and σe is a fixed edge bandwidth. The edge channel is then Iedge (u, v) = Σ(i,j)∈E Wij · max(xi, 0) · max(xj, 0)·gij (x)·segij (u, v), where √(xi xj) is the interaction intensity — non-zero only when both endpoint features are active for this sample — scaled by the relationship weight Wij. This channel is therefore a spatial map of which pairwise feature interactions are active for the specific row being encoded, distinct from the node channels, which encode only marginal feature values. The final image stacks all channels, I = [I1,..., IS, Iedge (, Igraph)], with C = S + 1 (or S + 2 with the graph channel); in our primary configuration (S = 2, edge channel included, no static graph channel) this gives C = 3.

TabSOM-derived interpretability analysis

To provide interpretability, we derive two methods from the trained SOM, which provide a two-dimensional representation of the joint feature distribution. Let the trained SOM consist of a grid of K = H × W nodes. Each node (r, c) has a prototype weight vector wr,c ∈ [0, 1]F, where F is the number of input features. The component plane of feature j is Φj (r, c) = wr,c,j, i.e. the j-th coordinate of the prototype at node (r, c). For each training sample xi, its best-matching unit BMU(xi) ∈ 1,..., H × 1,..., W is the node whose prototype is closest in Euclidean distance.

Self-organizing map class-separation importance: For binary classification, we define class-conditional activation densities over the SOM grid from the BMU assignments of the training set. For class k ∈ 0, 1, Ak (r, c) = i: yi = k, BMU(xi) = (r, c) / i: yi = k. A1 and A0 are normalized satisfying Σr,c Ak (r, c) = 1. The class-separation importance of feature j is the absolute difference between the A1- and A0-weighted averages of its component plane, Impj class-sep = Σr,c A1 (r, c) Φj (r, c) − Σr,c A0 (r, c) Φj (r, c). This quantity is large when feature j takes systematically different typical values in the map regions favored by each class. It is computed from the SOM together with the (binary) training labels, independent of downstream predictive models. We use Impj class-sep as the primary SOM-derived global feature-importance ranking.

Prototype-based partial dependence plot: Each SOM prototype wr,c represents a density-weighted combination of feature values drawn from the trained SOM’s organization of the input space. In contrast to the synthetic, independence-assuming Cartesian grids used by conventional Partial Dependence Plots (PDPs). We exploit this by encoding every prototype with the fitted encoder and passing the resulting image through the trained CNN fθ, yielding a prediction map Ψ(r, c) = fθ Encode(wr,c) ∈ [0, 1], interpreted as the predicted probability P (y = 1) for the (synthetic but data-consistent) sample represented by node (r, c). Ψ has the same spatial shape (H, W) as every component plane Φj. For a given feature j, we construct a prototype-based PDP by plotting the pairs Φj (r, c), Ψ(r, c) over all K nodes, where Φj (r, c) is mapped back to the feature’s original units via the inverse min-max transform. Each point is weighted by its hit count n(r, c) = i: BMU(xi) = (r, c), the number of training samples for which node (r, c) is the best-matching unit. A density-weighted mean curve is obtained by partitioning the nodes into B bins of approximately equal cumulative hit count, ordered by Φj, and computing the hit-count-weighted mean of Ψ within each bin. This curve is read analogously to a standard PDP — the expected model output as a function of feature j — but is restricted to regions of feature space the SOM actually populates with prototypes, so that sparsely-supported (low hit-count) segments are visually distinguishable from well-supported ones.

Results

Datasets: We evaluate TabSOM on four real-world datasets: Oxford Parkinson’s Disease (PAR), Pima Indians Diabetes (PID), QSAR Biodegradation (QSA), and Wisconsin Diagnostic Breast Cancer (WBC). All datasets are obtained from the public repository UCI Machine Learning Repository and present two classes (binary classification). The datasets have the following characteristics: PID (768 samples, 8 features), PAR (195 samples, 22 features), QSA (1055 samples, 41 features), WBC (569 samples, 30 features).

Experimental setup: We split each dataset into two independent subsets, a training subset (80% samples) and test subset (20% samples). To evaluate the generalization of predictive models, all methods are evaluated under 5-fold stratified cross-validation. Class imbalance was addressed through random undersampling, and it is applied only for training subset to prevent data leakage. All features are min-max normalized to [0, 1]. We evaluate the predictive performance using the Area Under the Receiver Operating Characteristic Curve (AUROC). Results are reported as the mean and standard deviation across five random seeds. TabSOM is configured as follows. The SOM grid side is set automatically to H = W = max(8, ⌈√(1.3F)⌉). Feature placement uses Hungarian assignment with centroid anchors, and the relational graph uses Pearson correlation between component planes with threshold τ = 0.5. The primary rendering configuration uses S = 2 node channels at σ ∈ 0.05, 0.08 and the relational edge channel. Encoded images are classified using a CNN with three convolutional blocks (16-32-64 channels), each followed by batch normalization and ReLU, with a max-pooling step after the second block — followed by global average pooling and a single linear output unit. The network is trained with Adam, the learning rate 10−3, weight decay 10−4, batch size of 32 and a class-balanced binary cross-entropy loss. The same architecture and optimizer are used for every tabular-to-image encoding method.

Classification results and benchmark: Table 3 compares TabSOM against twelve existing tabular-to-image encoding methods across six binary datasets. Across the four datasets, TabSOM ranks first in AUCROC on Pima (0.8236) and WDBC (0.9911), and third on PAR (0.8852) and QSAR (0.9098). These results yield the second-highest overall mean AUCROC (0.9024) and the second-best average rank, with only the Combination method performing better overall (mean AUCROC: 0.9114). The performance gap between TabSOM and the strongest competing method is small across all datasets. TabSOM outperforms the next-best method (Combination) by 0.0037 AUCROC on Pima and 0.0014 on WDBC, while underperforming Combination by 0.0327 AUCROC on Parkinsons and 0.0081 on QSAR. This places TabSOM and Combination as the two clearly strongest methods in the benchmark, consistently separated from the rest of the field by a substantial margin — the third-place method on most datasets (BarGraph or DistanceMatrix) trails TabSOM by roughly 0.01–0.04 AUCROC, while methods based on generic dimensionality-reduction embeddings (TINTO (tSNE), DeepInsight (tSNE), DeepInsight (UMAP), Fotomics) trail by considerably more, often falling 0.15–0.3 AUC below TabSOM and showing markedly higher variance (standard deviations frequently exceeding 0.10, compared to TabSOM’s 0.018–0.039 across all four datasets). Additionally, TabSOM’s standard deviation is among the lowest of any method on every dataset, indicating more stable performance across different seeds than most tabular-to-image approaches. Several methods (TINTO (tSNE and blur) on PAR, DeepInsight (tSNE) on PAR, Fotomics on WDBC) show standard deviations an order of magnitude larger, suggesting these embeddings are less reliable across different seeds. Finally, the gap between TabSOM and the weaker methods widens on datasets with fewer features (Pima, 8 features), where structured encodings (TabSOM, BarGraph, Combination, DistanceMatrix) outperform embedding-based methods by the largest margin (e.g., TabSOM exceeds TINTO (tSNE) by 0.30 AUCROC on Pima), while on higher-dimensional or more separable datasets (WDBC) most methods cluster more closely near the best performance.

Channel decomposition and feature maps: Figure 1 shows channel decomposition and final image resulting of the TabSOM encoding for samples from the PID. The sharp and mid node channels preserve each feature’s individual contribution at a stable canvas position across samples, with intensity tracking the feature’s normalized value. The edge channel shows the pairwise relationships captured by the correlation-based relational graph. A comparison across samples indicates that the overall structure of these channels is consistent among samples, whereas localized intensity differences indicate patient-specific feature values. Regarding feature relationships, for instance, sample 3 exhibits substantially stronger activation along the blood pressure–age and skin thickness–blood pressure edges than the other samples. This pattern is evident both in the edge channel and in the composite image, where it appears as a pronounced cyan–white region.

SOM-derived feature importance and dependence analysis: Figure 2 compares the feature ranking produced by SOM class separation against three established baseline importance measures: RF, XGB and SHAP. As shown, all four methods agree that glucose is the most important feature, and broadly agree in ranking BMI, age, and diabetes pedigree in the upper-middle tier while ranking blood pressure and skin thickness lowest. SOM class separation deviates most from the tree-based methods on age, for which it assigns the second-highest importance of any feature (0.54) versus a comparatively lower ranking from RF, XGB, and SHAP, and on pregnancies, for which it assigns higher relative importance than RF or SHAP. Despite these individual differences in ranking order, the overall agreement across methods derived from entirely different supports SOM class separation as a meaningful importance measure rather than an artifact of the encoding or grid placement procedure.

Figure 3 shows the prototype-based partial dependence curve for features of the PID. Four features, glucose, age, blood pressure, and insulin, exhibit increasing dependence, with predicted probability spanning roughly the full observed range (0.0 to 0.7–0.8). Glucose in particular shows an almost linear relationship, consistent with its established role as the most direct marker of glycaemic control in this domain. The remaining four features, pregnancies, bmi, skin thickness, and diabetes pedigree, show non-monotonic curves with a localized dip or spike in an overall increasing trend, and a visibly narrower range of predicted probability than the first group. diabetes pedigree shows the flattest and narrowest curve of all features (predicted probability varying only between approximately 0.12 and 0.52 in its observed range), consistent with its role as a comparatively weak indirect risk indicator relative to direct physiological measurements.

Discussion

In this paper, we proposed TabSOM, a tabular-to-image encoding method based on the SOM, and compared it against twelve state-of-the-art tabular-to-image methods across four benchmark datasets. TabSOM achieves competitive classification performance in all evaluated datasets, ranking first or second on every dataset while exhibiting the lowest variance of any method in the comparison. The benchmark comparison reveals two insights. Methods based on deterministic and spatially consistent encodings (BarGraph, Combination, DistanceMatrix, FeatureWrap, and BIE) perform similarly to TabSOM because they produce stable image representations. Methods based on stochastic dimensionality reduction, such as TINTO (t-SNE and PCA variants), DeepInsight (t-SNE and UMAP), IGTD, and Fotomics, perform substantially worse, particularly on smaller datasets such as Pima and QSAR, where the instability of the embedding across folds is reflected in larger standard deviations. TabSOM’s component-plane placement avoids this instability by deriving feature positions from the SOM geometry rather than from a dataset-level embedding, producing a layout that is fixed after a single SOM fit.

The interpretability analysis shows that TabSOM provides global explanations through the class-separation importance score and feature interactions via the prototype-based partial dependence. The feature ranking based on class-separation importance identifies features whose component planes are spatially separated by class label, features whose high or low values map systematically to different regions of the SOM grid depending on the outcome. In PID, glucose, age and BMI rank highest according to class-separation importance, consistent with their relevance as relevant factors associated with diabetes. The feature ranking comparison performed with Random Forest, XGBoost, SHAP, and SOM class-separation showed that the SOM-derived ranking agrees moderately with the model-based rankings, with the strongest agreement on the top-ranked features. This suggests that the SOM grid structure captures complementary information, providing interpretability of which features are globally important.

Additionally, the prototype-based partial dependence curves reveal the feature’s effect on the classifier’s output. For the PID, glucose showed a near-linear positive relationship across its observed range, consistent with its role as a glycaemic marker. Age exhibits a steep rise through the 30s and 40s before plateauing and diabetes pedigree produces the flattest and narrowest curve of all eight features, consistent with its status as a comparatively weak risk indicator. Several features (pregnancies, BMI, and skin thickness) showed localized non-monotonicities that the spatial overlay diagnostic traces to low-density regions of the SOM grid or to co-location with higher-ranked features, rather than to independent physiological effects.

In the literature, a single work has explored SOM in tabular-to-image methods. It encodes each sample as a proximity activation map over the SOM prototype grid, measuring Gaussian-weighted distances between the sample and every prototype in RBF kernel space, which produce images where pixel intensities reflect manifold position rather than individual feature values. TabSOM differs in three main dimensions: (i) it derives explicit per-feature canvas positions from SOM component planes via Hungarian assignment, producing spatially interpretable images in which each feature occupies a fixed region; (ii) it adds a relational edge channel encoding pairwise feature interactions not representable in any marginal-value image; and (iii) it provides interpretability tools (class-separation importance, prototype-based partial dependence) validated against established baselines, rather than qualitative activation-pattern inspection.

Beyond classification performance, interpretability analysis shows that the features identified by TabSOM align with established clinical knowledge, suggesting that TabSOM captures meaningful domain-relevant patterns. This property is especially valuable in high-stakes domains where transparency is as important as predictive performance. Future work will explore integration with multimodal data sources and systematic comparisons with established tabular interpretability methods. Future work will explore comparisons between Grad-CAM and established post-hoc methods for tabular data, such as SHAP, to better understand their strengths and limitations. Finally, future research may investigate the use of alternative architectures, such as vision transformers, to further enhance the representation learning capabilities of TabSOM-generated images.

Conclusions

This paper proposed TabSOM, a tabular-to-image encoding that uses the SOM to provide both a topology-based feature placement and a relational graph over features, rendered as a multi-channel image that separates marginal feature values from pairwise interactions. We benchmarked against twelve existing tabular-to-image methods on four public datasets. TabSOM ranked first and second on every dataset and exhibited the lowest standard deviation across all methods. TabSOM provides two interpretability tools, including a prototype-inspired partial dependence plot, and the class-separation importance score, to provide feature importance and ranking, and the dependence analysis. A comparative analysis of feature importance with SHAP, and RF showed agreement on the top-ranked features, consistent with the SOM grid capturing structural patterns beyond those encoded by tree-based impurity measures.

Improvements for AI systems

Improvements to AI systems based on TabSOM:

  1. Add relational feature encoding to tabular deep learning models. Most current tabular DL approaches (e.g., FT-Transformer, TabNet) treat features as independent or rely on attention mechanisms that do not explicitly encode pairwise feature interactions. TabSOM’s relational edge channel provides a principled way to inject pairwise feature co-variation into the input representation, enabling CNNs/ViTs to exploit feature interactions that are otherwise invisible to the model. An improved system could automatically learn which feature pairs matter for a given task, rather than relying on hand-crafted interaction terms.

  2. Replace stochastic dimensionality-reduction layouts with deterministic, topology-preserving placements. Existing tabular-to-image methods (t-SNE, UMAP, PCA-based) produce unstable layouts across random seeds, causing high variance in downstream model performance. TabSOM’s SOM-based placement with Hungarian assignment yields a fixed, reproducible canvas after a single fit. An improved system could guarantee stable image representations across runs, reducing variance in model predictions and making results more reproducible for clinical or financial applications.

  3. Enable multi-scale feature rendering without manual bandwidth tuning. TabSOM stacks multiple Gaussian bandwidths as aligned channels, allowing a CNN to jointly use sharp (fine-grained) and smooth (regional) information. An improved system could automatically adapt the number and range of bandwidths based on feature density and dataset size, eliminating the need for users to choose a single smoothing parameter and improving performance on datasets with heterogeneous feature distributions.

  4. Provide built-in, model-agnostic interpretability from the encoding itself. TabSOM’s class-separation importance score and prototype-based partial dependence plots derive explanations directly from the SOM grid, independent of the downstream classifier. An improved AI system could offer interpretability without requiring post-hoc tools like SHAP or LIME, which are computationally expensive and can be unstable. This is especially valuable in high-stakes domains (healthcare, finance) where model transparency is mandatory.

  5. Generate data-consistent partial dependence curves. Unlike standard PDPs that assume feature independence and sample from synthetic Cartesian grids, TabSOM’s prototype-based PDP uses only regions of feature space actually populated by the SOM’s prototypes, weighted by hit counts. An improved system could produce more trustworthy dependence analyses that avoid extrapolating to unrealistic feature combinations, reducing misleading insights in scientific or medical decision-making.

  6. Create a unified pipeline for feature placement, interaction discovery, and explanation. TabSOM integrates layout, relational graph, and interpretability into a single SOM-based framework. An improved AI system could offer an end-to-end solution where the same learned representation serves both prediction and explanation, eliminating the need to maintain separate models for classification and interpretability, and reducing deployment complexity.

  7. Improve performance on small and low-dimensional tabular datasets. TabSOM’s structured encoding significantly outperforms embedding-based methods on datasets with few features (e.g., 8-feature Pima dataset, 0.30 AUCROC improvement over TINTO). An improved system could leverage this approach for small-sample medical or sensor datasets where deep learning typically underperforms, enabling robust DL application in data-scarce domains.

  8. Support class-conditional feature importance for binary and multi-class problems. TabSOM’s class-separation score quantifies how differently a feature’s component plane is activated across classes. An improved system could extend this to multi-class settings and provide per-class feature rankings, helping users understand which features drive each class decision—useful for fraud detection (distinguishing fraud types) or medical diagnosis (differentiating disease subtypes).

  9. Enable visual inspection of feature assignments and class separability. TabSOM’s feature maps allow users to see where each feature is placed and how well classes separate on the SOM grid. An improved AI system could provide interactive visualizations that let practitioners verify that the encoding preserves meaningful structure before training a classifier, improving trust and debugging capability.

  10. Reduce computational overhead compared to stochastic embedding methods. TabSOM requires only a single SOM fit (unsupervised) plus a Hungarian assignment, which is deterministic and fast. An improved system could avoid repeated t-SNE/UMAP runs across cross-validation folds, cutting preprocessing time and enabling real-time or near-real-time tabular-to-image conversion for streaming or online learning applications.

Abstract

Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction method (e.g., t-SNE, UMAP, PCA). However, they encode only the marginal value of each feature and discard information about feature relationships. We propose TabSOM, a tabular-to-image encoding built on the Self-Organizing Map (SOM), which provides: (i) a spatial layout in which every input feature occupies a fixed canvas position derived from its component plane via collision-free Hungarian assignment; and (ii) a graph that captures pairwise feature relationships derived from the SOM component planes. The resulting image stacks two multi-scale node channels: one encodes feature values at fixed scales, while the other encodes pairwise feature interactions as spatial connections between related features. Two SOM-derived interpretability approaches are introduced: a prototype-inspired partial dependence plot and a class--separation importance score. Benchmarked against twelve existing tabular-to-image methods across public binary-classification datasets, TabSOM ranks first or second on every dataset and achieves the lowest variance of any method evaluated. Interpretability obtained with TabSOM was validated against Random Forest, XGBoost, and SHAP, the class-separation score shows reasonable agreement with established baselines on the top-ranked features while capturing complementary structural information from input data. These results demonstrate that TabSOM provides an effective and interpretable approach for applying deep learning architectures to tabular data, bridging the performance--interpretability gap in this domain.

Related papers