Bloom Filter Encoding for Machine Learning

arXiv:2512.19991 · cs.LG · Submitted 2025-12-23 · Read on arXiv

cs.LG

Submitted: 2025-12-23

Updated: 2026-05-08

Comments: 14 pages, 7 figures

Journal ref: Artificial Intelligence Applications and Innovations (AIAI 2026), IFIP Advances in Information and Communication Technology, vol. 792, pp. 17-31, Springer, 2027

DOI: 10.1007/978-3-032-30612-8_2

License: http://creativecommons.org/licenses/by/4.0/

The gist: We present a method that uses a Bloom filter transform to preprocess data for machine learning.

Terminology

Abstract

We present a method that uses a Bloom filter transform to preprocess data for machine learning. Each sample is encoded into a compact bit-array representation using hash-based encoding, producing a fixed-length feature space that reduces memory usage and obfuscates original feature values. The encoding does not rely on keyed hashing; however, a key can optionally be used to control the mapping and would be required to reproduce the representation. We evaluate the approach on six datasets spanning text, time-series, tabular, and image domains: SMS Spam Collection, ECG200, Adult 50K, CDC Diabetes, MNIST, and Fashion MNIST. Four classifiers are considered: Extreme Gradient Boosting, Deep Neural Networks, Convolutional Neural Networks, and Logistic Regression. Results show that models trained on Bloom filter encodings achieve performance comparable to models trained on raw data or standard dimensionality reduction techniques across several datasets, while providing consistent memory savings. These findings suggest that Bloom filter encodings can serve as an efficient, general-purpose pre-processing representation that preserves useful similarity structure for learning tasks while providing a degree of data obfuscation.

Related papers