AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "AI Appeals Processor".
Jane: The gist The AI Appeals Processor presents a microservice-based system that integrates natural language processing and deep learning techniques for automated classification and routing of citizen appeals,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we’re looking at "AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services" and the authors are Vladimir Beskorovainyi Besk from MIPT, who built this whole system to automate how agencies sort appeals. They point out that the old way takes about twenty minutes per appeal with only sixty-seven percent accuracy, which is a major slowdown for public service.
Jane: The core idea here is integrating natural language processing and deep learning techniques into a microservice-based system specifically for classifying and routing citizen appeals, tackling the complexity of these submissions electronically now that they are increasing significantly.
Lu: What’s interesting is how they immediately evaluate several different approaches, starting with simpler methods like Bag-of-Words with SVM and TF-IDF with SVM before settling on their deep learning solution for this Russian language context.
Meng: And the paper introduces a set of five principal components for this system, which includes a user interface, a backend server, an AI module written in Python with TensorFlow, a database for storage, and an integration layer that connects to external document management systems. That’s quite a comprehensive setup to start with.
Lalam: It suggests that by using this microservice approach, they can make sure every part of the system is independently developed and scalable, which is important when you’re dealing with the high volumes of text these government agencies get.
Tom: So what we found in "AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services" is that their Word2Vec plus LSTM architecture achieved seventy-eight percent classification accuracy on a dataset of ten thousand real citizen appeals.
Jane: That seventy-eight percent accuracy is quite respectable for Russian language appeals, and they also showed a significant speed improvement, reducing processing time by fifty-four percent, dropping it from twenty-two point five minutes down to just ten point two five minutes per appeal.
Lu: That speed reduction is substantial because it directly addresses the bottleneck they mentioned earlier in the paper; getting twenty-two point five minutes down to around ten point two five means a much faster service for citizens and staff, which is what they were aiming for when they looked at this problem.
Meng: But they also gave us a breakdown of how well this model performed on different appeal types, showing that Proposals were classified most accurately at eighty point eight percent, followed by Complaints at seventy-eight point one percent and Applications at seventy-five point nine percent.
Lalam: It really demonstrates that combining word embeddings with a recurrent network is effective for capturing the necessary meaning for sorting while keeping the system efficient enough to handle a large amount of text.
The paper's summary: Tom: So now let’s talk about what they actually found in "AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services." The core finding is that their Word2Vec plus LSTM architecture achieved a seventy-eight percent classification accuracy when tested on the dataset of ten thousand real citizen appeals.
Jane: That seventy-eight percent accuracy is quite respectable for Russian language appeals, and they also showed a significant speed improvement, reducing processing time by fifty-four percent, dropping it from twenty-two point five minutes down to just ten point two five minutes per appeal.
Lu: That speed reduction is substantial because it directly addresses the bottleneck they mentioned earlier in the paper; getting twenty-two point five minutes down to around ten point two five means a much faster service for citizens and staff, which is what they were aiming for when they looked at this problem.
Meng: But they also gave us a breakdown of how well this model performed on different appeal types, showing that Proposals were classified most accurately at eighty point eight percent, followed by Complaints at seventy-eight point one percent and Applications at seventy-five point nine percent.
Lalam: It really demonstrates that combining word embeddings with a recurrent network is effective for capturing the necessary meaning for sorting while keeping the system efficient enough to handle a large amount of text.
The paper's improvements: Tom: The authors don't just stop at that seventy-eight percent accuracy, they actually outline several ways they think this system could be made even better. They focus on how they can improve accuracy and efficiency further by changing the classification method itself.
Jane: One of the main suggestions is moving toward a multi-label classification approach to handle dual-intent appeals, which are those tricky situations where a citizen writes text that contains both a complaint and a request at the same time.
Lu: That makes sense because they noted that dual-intent appeals accounted for forty-two percent of their initial classification mistakes, so trying to force one category when there are two distinct needs is going to lead to errors.
Meng: They also suggest integrating a more advanced language model, specifically something like RuBERT, which is pre-trained for Russian instead of just relying on Word2Vec embeddings that they trained only on the ten thousand sample corpus.
Lalam: Using a pre-trained Russian model should really help the system handle those ambiguous appeals better because it understands the language structure more deeply than just looking at word vectors learned from a small set.
Conclusion: Tom: So to wrap up on "AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services," the main point is that this Word2Vec plus LSTM setup gives you seventy-eight percent accuracy and cuts processing time by fifty-four percent compared to the old manual ways.
Jane: It shows that for complex, real-world tasks involving Russian language text, combining embeddings with recurrent networks is a very effective way to build a usable AI system for handling high volumes of information.
Lu: The path forward seems clear; they point toward using transformer-based models with Russian pretraining like RuBERT to tackle those trickier multi-intent appeals and boost performance even higher.
Meng: From an engineering standpoint, the architecture itself—that microservice design—is what makes it scalable enough to handle twenty concurrent requests without a massive slowdown in the actual running time.
Lalam: It really changes how we think about public service automation; it moves us from slow, manual review to a system that can process appeals much faster and with much better initial accuracy.
Tom: The limitation they mention is that the Word2Vec embeddings were trained on this limited corpus instead of using pre-trained Russian word vectors like fastText or navec.
Jane: That means the quality of those initial representations might be capped by how much data they started with, so future work needs to focus on richer language foundations.
Lu: Exactly; if you feed a model a small dataset, it's only as good as what that small dataset can teach it about the language structure overall.
Meng: For practical application, that means if we want this system to handle even more specialized domains, we might need to feed it a much bigger library of pre-existing Russian language knowledge.
Lalam: It really changes how we think about public service automation; it moves us from slow, manual review to a system that can process appeals much faster and with much better initial accuracy.
Tom: So while this Word2Vec plus LSTM setup is solid for a baseline, the real potential lies in integrating those more advanced transformer models they mentioned.
Jane: It shows that the immediate win is efficiency—that fifty-four percent time reduction—but the long-term gain comes from using better language models to get that accuracy even higher.
Lu: I think we should also keep an eye on how they suggest expanding the classification taxonomy beyond those seven primary domains because those specific, technical areas like housing and utilities are where the real complexity lies.
Meng: That granularity is important for engineering; if you can't sort it into a clear category, you can't actually automate the routing or response process effectively.
Lalam: It really changes how we think about public service automation; it moves us from slow, manual review to a system that can process appeals much faster and with much better initial accuracy.
Tom: That's what we’re looking at in "AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services." It’s a great starting point for how deep learning can handle real-world, specific language tasks.
Jane: And if you want to see how those more advanced models are doing, we’ve got another paper coming up about federated learning and detecting when clients start acting like free riders.
Besk Tech · Moscow Institute of Physics and Technology (MIPT)
cs.CL, cs.AI
Submitted: 2026-04-04
Updated: 2026-10-08
Comments: 11 pages, 6 tables. v2 corrects the reference list, adds confidence intervals, a description of the production audit, and sections on lessons from deployment, limitations and ethical considerations; all test-set results are unchanged
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: The gist The AI Appeals Processor presents a microservice-based system that integrates natural language processing and deep learning techniques for automated classification and routing of citizen
Key concepts
- Word2Vec+LSTM
- This is the specific deep learning architecture used for text classification. Word2Vec creates numerical representations (embeddings) of words, capturing their meaning in context. The LSTM (Long Short-Term Memory) network then processes these word meanings sequentially to accurately classify the appeal type.
- Text Processing Pipeline
- This is the step-by-step process applied to raw text before classification. It involves cleaning the text (removing stop words, normalizing characters), turning words into numbers (feature extraction), and finally feeding those numbers into the neural network for a decision.
- Human-in-the-Loop
- This refers to a system where human operators review and verify the AI's classifications. When humans correct the model's mistakes, this feedback is used to retrain and improve the AI further. This cycle helps boost accuracy beyond its initial training level.
- Microservice-based System
- The system is built as a collection of small, independent services rather than one large program. Each part—like the frontend, backend server, or AI module—runs separately but communicates to perform the overall task. This design makes the system flexible and easier to update.
Terminology
Summary
The gist The AI Appeals Processor presents a microservice-based system that integrates natural language processing and deep learning techniques for automated classification and routing of citizen appeals, achieving 78% classification accuracy with Word2Vec+LSTM while reducing processing time by 54% >
System Overview
The AI Appeals Processor is implemented as a microservice-based web application comprising five principal components: (1) a user-facing frontend built with a JavaScript framework, (2) a backend server implemented in an object-oriented server-side framework, (3) an AI module for text processing and classification developed in Python with TensorFlow, (4) a relational database for persistent storage, and (5) an integration layer providing REST API connectivity to external document management systems >
Text Processing Pipeline
The text processing pipeline consists of four stages: Preprocessing, Feature extraction, Classification, and Post-processing >
-
Preprocessing involves normalization (lowercasing, special character removal), tokenization, stop-word removal using NLTK’s Russian stop-word list, and lemmatization via pymorphy2 >
-
Feature extraction converts preprocessed tokens to numerical representations using either TFIDF vectorization or Word2Vec embeddings trained on the appeal corpus >
-
Classification passes these extracted features to a classification model, with the production configuration using an LSTM-based neural network >
Model Evaluation and Selection
The paper evaluated six approaches for Russian-language citizen appeal classification on a dataset of 10,000 appeals > The models evaluated included: Baseline (Manual classification by human operators), BoW + SVM, TF-IDF + SVM, fastText, Word2Vec + LSTM, and BERT > The Word2Vec+LSTM architecture was selected for production deployment as it provides the optimal balance between classification quality (78% accuracy) and computational requirements > This model achieved 78% classification accuracy while reducing processing time by 54% (from 22.5 to 10.25 minutes per appeal) >
Performance Benchmarks
The Word2Vec+LSTM model achieved a classification performance of 78% accuracy on the test set > Per-class analysis showed that Proposals were classified most accurately (80.8%), followed by Complaints (78.1%) and Applications (75.9%) > Domain-specific classification demonstrated that Housing & Utilities achieved an 86% precision, recall, and F1-score > The system achieves a consistent 53–56% reduction in processing time across all appeal lengths compared to manual workflows > Under concurrent load of 20 requests, the processing time increased by only 25% compared to the single-request baseline >
Human-in-the-Loop and Scalability
The system incorporates a human-in-the-loop feedback cycle where iterative retraining on operator-verified data improved classification accuracy beyond the initial 78% achieved on the research dataset > Stress testing confirmed stable performance under concurrent loads, with processing time increasing only to 12.80 minutes at 20 concurrent requests > The system maintains acceptable performance under concurrent loads of up to 100 simultaneous requests, with model inference alone averaging under 2 seconds per appeal regardless of concurrency >
Comparative Analysis and Limitations
The proposed system achieves higher accuracy (78%) than existing solutions, which typically achieve classification accuracy of 45–65% [Wirtz et al.
Improvements for AI systems
-
Implement multi-label classification for appeals: Address
dual-intent appeals
by adopting a multi-label approach to classify multiple appeal types simultaneously, as suggested in the error analysis section:(1) dual-intent appeals (42% of errors), where citizens express both a complaint and a request within the same text.
-
Integrate RuBERT for superior language modeling: Replace the
fine-tuned bert-base-multilingual-cased
with a Russian-specific pretraining model like RuBERT to improve performance on ambiguous appeals, as suggested in the future work section:incorporating transformer-based models with Russian-specific pretraining (e.g., RuBERT) to further improve performance on ambiguous appeals.
-
Enhance domain classification granularity: Expand the classification taxonomy beyond the seven primary domains by implementing a hierarchical or multi-label structure to better capture nuanced citizen needs, as indicated by the finding that
domains with specialized terminology (housing and utilities, healthcare) are inherently more amenable to automated classification than heterogeneous categories.
-
Utilize pre-trained Russian word vectors for Word2Vec: Improve the quality of learned representations by training Word2Vec embeddings on pre-trained Russian word vectors (e.g., fastText or navec) instead of solely on the limited 10,000-sample corpus, mitigating the constraint noted in
Limitations
:the Word2Vec embeddings were trained on this limited corpus rather than using pre-trained Russian word vectors (e.g., fastText or navec).
-
Automate full processing pipeline: Transition from an end-to-end time that includes human review to a fully autonomous system by eliminating the
human-in-the-loop
step for classification, as implied by the goal of reducingfully autonomous processing times would be substantially lower.
Abstract
Government agencies must register, classify and route every citizen appeal within statutory time limits, and much of this work is still done by hand. We describe AI Appeals Processor, a classification and routing component deployed in a CPU-only government environment, and report what its evaluation and deployment taught us. On 10,000 real Russian-language appeals from a cross-domain dataset, we compare Bag-of-Words and TF-IDF with SVM, fastText, Word2Vec+LSTM and multilingual BERT on a three-way appeal-type task. On a held-out test set of 1,500 appeals, BERT reaches 82% accuracy and Word2Vec+LSTM 78%, against 67% for individual operators measured on an expert-adjudicated gold standard. We deployed the LSTM: in a workflow where an operator verifies every prediction, its lower training cost made frequent retraining on operator-verified labels practical, while the four-point accuracy gap did not change the operator's task. End-to-end handling time fell by 53-56% across four appeal-length bands (unweighted mean 22.5 to 10.25 minutes); model inference takes under two seconds of this. Most residual errors trace to the label taxonomy rather than the model: the statutory definitions of complaints and applications overlap, and many appeals carry two intents. A post-deployment audit of production classifications, made after several retraining cycles by operators who saw the assigned category, judged more than 95% correct; we explain why this figure is not comparable with the test-set result.
Sources
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering