FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees

arXiv:2608.12140 · cs.AR, cs.LG · Submitted 2026-08-12 · Read on arXiv

Zhiqiang Que, Chang Sun, Haiyang Wang, Dinesh Pamunuwa, Roshan Weerasekera, Qijia Tang, Bakhtiar Zadeh, Wayne Luk, Maria Spiropulu

University of Bristol · California Institute of Technology · Imperial College London

cs.AR, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: accepted by ASAP'26. Code available at https://github.com/ecs-bristol/FQTree

Code: https://github.com/ecs-bristol/FQTree

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees This paper presents the FQTree algorithm for fine-grained quantization-aware training of boosted decision trees

Terminology

Summary

FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees

This paper presents the FQTree algorithm for fine-grained quantization-aware training of boosted decision trees (BDTs), together with the QXGB framework for automatic hardware generation targeting low-latency FPGA deployment. The work addresses the challenge that existing FPGA-based BDT implementations typically rely on uniform or manually tuned fixed-point formats, which can introduce unnecessary hardware cost or accuracy loss.

The main contributions are: (1) the FQTree algorithm, a hardware-aware and fine-grained quantization-aware training approach featuring a hardware-oriented leaf-value quantization formulation with a global quantization step, tree-wise shift, and bias folding; (2) the QXGB framework, a compiler-based scalable flow for generating efficient low-latency FPGA-based BDT implementations for both high-level synthesis (HLS) and RTL design; and (3) a comprehensive evaluation showing LUT usage reduction of 26–57% compared to state-of-the-art BDT designs while matching or improving accuracy and achieving lower latency.

The methodology introduces several key innovations. For leaf-value quantization, instead of assigning one uniform fixed-point format to the whole ensemble, FQTree quantizes each tree's leaf values according to their own numerical range. The quantization is defined as: v' = round(max(s·(v − f(v)), 0)), v qint = v' − min(v'), v qtrain = (v qint + min(v'))·s−1 + f(v), where v qint is the integer representation used for hardware generation, s is the global step size, and f(v) is a tree-wise shift factor. The maximum effective bitwidth of one tree is determined by bt = ⌈log2(max(v qint) + 1)⌉, so trees with smaller post-quantization ranges naturally require fewer bits in hardware.

The shift factor f(v) serves two purposes: it pushes leaf values toward a non-negative representation to reduce selection logic cost, and it enables controlled clipping/pruning by increasing zeros in the quantized leaf vector. The offset removed during quantization is folded into a tree-level or class-level bias term, added only once after tree accumulation, which is much cheaper than carrying signed offsets through every tree path.

A key aspect of the approach is that quantization is applied during boosting, so later trees adapt to the errors of the already-quantized ensemble rather than an ideal floating-point one. This allows subsequent boosting stages to compensate not only for task residuals but also for quantization distortion. The paper notes that the earliest boosting rounds require substantially more bits, while later rounds gradually become less demanding in terms of leaf-value precision, consistent with the boosting process where early trees capture dominant decision patterns.

For hardware generation, the QXGB framework extends the DAIS intermediate representation with an explicit MUX instruction to capture conditional selection behavior of decision nodes. The compiler lowers each internal decision node to a fixed-point comparison between quantized input feature and quantized threshold, transforming each tree into a comparator–mux network. The generated accelerator consists of four stages: input quantization/alignment, node evaluation with fixed-point comparisons, tree evaluation using MUX hierarchies, and ensemble accumulation using an adder tree where heterogeneous leaf-value quantization leads to narrower accumulators locally.

The evaluation covers three datasets: MNIST handwritten digit classification, the OpenML jet substructure classification (JSC) dataset in high-energy physics, and a binarized version of the UNSW-NB15 network intrusion detection (NID) dataset. The target device is xcvu13p-flga2577-2-e, using Vivado 2025.1 for resource and Fmax measurements, with Verilator for bit-and-cycle-accurate verification.

Key results include: On JSC HLF, FQTree achieves 75.7% accuracy with 1,652 LUTs and 2 cycles (4.0 ns) latency, compared to TreeLUT's 75.6% accuracy with 2,234 LUTs and 3 cycles, representing about 26% LUT reduction. At a lower-cost point, FQTree achieves 74.8% accuracy using only 548 LUTs and 1 cycle (2.0 ns), compared to TreeLUT's 74.6% accuracy, 796 LUTs, and 2 cycles, corresponding to about 31% fewer LUTs. On MNIST, FQTree's highest-accuracy configuration reaches 97.7% with 8,147 LUTs and 2 cycles, outperforming the best previously reported accuracy of 97.2% from POLYBiNN while requiring far fewer LUTs. At a moderate point, FQTree achieves 96.7% accuracy with 2,744 LUTs versus TreeLUT's 96.6% with 4,478 LUTs, giving about 39% LUT reduction. On NID, FQTree achieves 93.1% accuracy with only 157 LUTs and 1 cycle, compared to TreeLUT's 92.7% accuracy, 345 LUTs, and 2 cycles, corresponding to about 55% LUT reduction.

The paper also compares against post-training quantization (PTQ) baselines using the same quantizer, showing that FQTree improves over PTQ by incorporating quantization during training. The design space exploration shows that FQTree exposes a rich set of implementation points spanning different resource budgets, with moderate tree depths (especially max depth=4 for JSC) providing the most favorable accuracy–resource trade-offs. The paper concludes that FQTree provides a practical way to explore and select quantized BDT implementations matching different application-level accuracy targets and FPGA resource budgets, with future work extending to larger BDTs, broader FPGA platforms, and trustworthiness-aware BDT inference.

Improvements for AI systems

Improvements to AI systems:

  1. Quantization-aware boosting for tree ensembles – Integrate FQTree’s fine-grained, per-tree leaf-value quantization directly into the training loop of any gradient-boosted decision tree (GBDT) library (e.g., XGBoost, LightGBM). This allows the model to adapt to quantization errors during boosting, reducing accuracy loss compared to post-training quantization.

  2. Hardware-aware model selection – Use FQTree’s design space exploration to automatically generate a Pareto frontier of accuracy vs. FPGA resource usage (LUTs, latency) for a given dataset. This enables an AI system to pick the optimal BDT configuration (tree depth, number of trees, quantization bits) for a target hardware budget without manual tuning.

  3. Automatic FPGA accelerator generation – Extend QXGB’s compiler flow to produce both HLS and RTL implementations from a trained BDT, with heterogeneous leaf-value bitwidths and tree-wise shift/bias folding. This reduces LUT usage by 26–57% and latency by 1–2 cycles compared to state-of-the-art designs, enabling deployment on smaller or cheaper FPGAs.

  4. Adaptive precision scheduling during training – Leverage the finding that early boosting rounds need more bits than later ones to implement a dynamic bitwidth allocation strategy. This can reduce memory and compute during training while maintaining accuracy, and also produce more hardware-efficient models.

  5. Edge AI inference with ultra-low latency – Deploy the generated accelerators for real-time applications (e.g., network intrusion detection, jet substructure classification) achieving 1–2 cycle latency (2–4 ns) and sub-200 LUT usage for simple tasks, enabling AI inference on resource-constrained edge devices.

  6. Trustworthiness-aware BDT inference – Extend the framework to incorporate robustness metrics (e.g., adversarial robustness, calibration) during quantization-aware training, allowing the generated hardware to maintain reliability under input perturbations.

What the improved AI system can do:

  • Train BDTs that are natively optimized for FPGA deployment, achieving state-of-the-art accuracy with 26–57% fewer logic resources and lower latency than existing methods.

  • Automatically generate a family of hardware implementations (from ultra-low-resource to high-accuracy) for a given task, allowing system designers to trade off accuracy vs. cost in seconds.

  • Deploy real-time inference on FPGAs for high-throughput applications (e.g., network packet classification at line rate, particle physics event filtering) with deterministic, cycle-accurate behavior.

  • Support heterogeneous quantization across trees, reducing memory footprint and power consumption without sacrificing model quality.

  • Provide a scalable path from algorithm to silicon, reducing design time for custom AI accelerators by replacing manual fixed-point tuning with automated, training-aware quantization.

Abstract

Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or manually tuned fixed-point formats, which can introduce unnecessary hardware cost or accuracy loss. This work presents the FQTree algorithm https://github.com/ecs-bristol/FQTree for fine-grained quantization-aware training of BDTs, together with the QXGB framework for automatic hardware generation. FQTree introduces a hardware-oriented leaf-value quantization scheme that uses a global quantization step together with a tree-wise shift, enabling compact non-negative integer leaf representations, controlled clipping/pruning, and bias folding to reduce datapath cost. This work further applies this quantization during boosting so that later trees adapt to the errors of the already-quantized ensemble, and then lowers the trained model into low-latency hardware implementations through a compiler-based flow. Results on JSC, MNIST, and NID show that our method reduces LUT usage by 26-57% compared with the state-of-the-art FPGA-based BDT designs while matching or improving accuracy.

Related papers