LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

arXiv:2607.15509 · cs.CV, cs.AI, cs.LG · Submitted 2026-08-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "LLM-Driven AutoML for Cross-Lingual Handwritten OCR".

Jane: The paper, "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re looking at the paper titled "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-five GPT-4o, and Claude Sonnet four" and the authors are Mobina Kashaniyan Amirhossein Ghassemi from Iran University of Science and Technology. It sounds like this paper is tackling the big problem of handwritten text recognition across different scripts using some very advanced language models.

Jane: That’s right, Tom, and it really highlights how these large language models can move beyond just generating text to becoming active participants in building complex machine learning systems for tasks like OCR. It suggests a new way of thinking about designing these kinds of tools.

Lu: I find the use of GPT-five GPT-4o, and Claude Sonnet four as independent design agents really interesting; it opens up possibilities for exploring architectures we might never even consider manually because human intuition has limits in that space.

Meng: From an engineering standpoint, the authors are focusing on creating a fully automatic pipeline that handles everything from data preparation to model training without needing constant human input for every step. That’s a big operational win if you’re trying to scale up these kinds of recognition systems.

Lalam: The title itself points to this "Closed-Loop Neural Architecture Search," which implies the system doesn't just build one thing and stop; it keeps testing and improving the design based on its own performance results. That kind of continuous self-improvement feels very promising for cultural preservation tasks, which is what this whole project is focused on.

Tom: Exactly, Jane; that closed loop means the system gets smarter with every trial, which moves it away from static designs we use today. This paper is showing how to automate that iterative refinement process using the power of these large language models. Where do we go from here?

The paper's summary: Jane: Now that we know the setup, let’s look at what the paper actually describes in terms of its core methodology. Essentially, they are taking raw handwritten images from different sources like English with EMNIST, Persian with SADRI, and Arabic with AHCD.

Tom: And what they do is feed those images into a pipeline where the AI models—GPT-five through Claude Sonnet four—are tasked with proposing the actual neural network structures for recognition. It’s not just guessing; it’s generating specifications in a structured format, like JSON, which then get converted into working Keras models.

Lu: What I find particularly compelling is how they tailor the preprocessing and encoding strategies specifically for each script to match its unique visual complexities, which addresses the initial problem statement about different languages demanding unique solutions.

Meng: They also mention using a lightweight data augmentation strategy, like random rotations and shifts, during training with Keras’s ImageDataGenerator to help the model generalize better without adding too much computational load during the search phase. That shows they are thinking about practical deployment constraints right from the start.

Lalam: This whole summary points to a system that is cross-lingual by design, meaning it’s not just a set of separate tools for English and Arabic, but one unified framework that handles them all within the same automated search process.

Tom: So, they are using these LLMs as the primary designers to propose architectures that fit the specific data they feed them, which is a major departure from how we typically design neural networks today. Jane, what does this mean for practical implementation?

The paper's improvements: Jane: The main improvement they highlight is completely removing the need for human intervention in model selection and hyperparameter tuning; instead, the system handles that entire iterative optimization cycle automatically. This eliminates a major bottleneck that used to slow down development significantly in traditional OCR work.

Tom: That’s a huge operational advantage, Jane; being able to scale up across many languages without needing a team of experts for every single model design is something we need to focus on for real-world adoption. Lu, what about the performance aspects?

Lu: The authors show that the system leverages optimal architectural patterns and hyperparameters specific to each targeted handwriting script, which means they aren't just randomly picking structures; they are using knowledge of how those scripts are visually organized.

Meng: They also find that this automated process maintains inference speeds that meet the needs of real-time applications, and they noted that the system explores a wide range of neural architectures adaptively based on what it learns during the search. That adaptation is key for efficient deployment.

Lalam: This capability to automatically select designs that fit the unique requirements of each script, without explicit human guidance, really speaks to how adaptable these AI systems can be when applied to diverse cultural data sets.

Tom: It seems like they've moved the focus from designing a single perfect model to building a scalable system capable of discovering many efficient models tailored for many different languages. This paper is showing a new path for AI development in this area.

Conclusion: Jane: So, to wrap up our discussion on "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-five GPT-4o, and Claude Sonnet four" the core implication is that these LLMs are acting as autonomous generators and evaluators rather than just passive tools.

Tom: Exactly, Jane; it's about a shift from relying on human intuition for design to having an autonomous system discover and refine the best architectures through a closed-loop process. This whole paper shows how effective this approach can be when applied across multiple languages simultaneously.

Lu: The potential for these models to discover novel architectures is something we will definitely see expanded upon in other areas, perhaps in analyzing complex visual patterns far beyond just text recognition.

Meng: For my team, the practical implication is that this gives us a framework for building systems that adapt to real-world handwriting complexity without being constrained by pre-defined designs or needing constant manual tuning cycles from human experts.

Lalam: We can use this to ensure that historical records aren't lost because they're too complex or expensive to digitize using conventional methods, making knowledge more accessible globally through this technology.

Tom: This entire journey from manual model building to seeing "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-five GPT-4o, and Claude Sonnet four" is a significant step forward in how we do AI research.

Jane: It’s clear that the system is achieving both robust performance across Arabic, English, and Persian while maintaining good efficiency metrics which makes this a very compelling piece of research.

Lu: I think the potential for finding optimal patterns across different language structures suggests that the very definition of how we design AI systems is changing here.

Meng: It provides a clear pathway to rapid and adaptable deployment of OCR technology across many languages and domains, addressing real-world operational needs that were previously unmet by slower methods.

Lalam: We can use this to ensure that historical records aren't lost because they're too complex or expensive to digitize using automated means.

Tom: Thanks everyone for joining me on the show; we’ll be right back with the next paper from arXiv!

Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani

Iran University of Science and Technology Department of Computer Engineering, Tehran, Iran · Iran University of Science and Technology Department of Computer Engineering, Tehran, Iran

cs.CV, cs.AI, cs.LG

Submitted: 2026-08-18

Updated: 2026-08-20

Comments: 7 pages, 10 figures, and 2 tables. Published in the 2025 15th International Conference on Computer and Knowledge Engineering (ICCKE), IEEE

Journal ref: 2025 15th International Conference on Computer and Knowledge Engineering (ICCKE), IEEE, 2025, Article No. 11273810

DOI: 10.1109/ICCKE68588.2025.11273810

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 80/100

The gist: The paper, "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4," addresses the enduring challenge of handwritten text

Key concepts

LLM-Driven AutoML
This involves using large language models like GPT to act as design agents that propose neural network structures for OCR tasks. These models generate specifications, such as JSON, which are then converted into working Keras models, automating the design process.
Closed-Loop Neural Architecture Search
This methodology means the system continuously tests and improves its own model designs based on performance results. It does not stop after one design but iteratively refines the architecture to find optimal structures for recognition.
Cross-Lingual Handwritten OCR
This refers to a system that can recognize handwritten text from different scripts, specifically English (using EMNIST), Persian (using SADRI), and Arabic (using AHCD). The paper focuses on creating a unified framework for these diverse languages.
Autonomous Design Agents
The use of multiple large language models—GPT-5, GPT-4o, and Claude Sonnet 4—as independent design agents allows the system to explore model architectures that human intuition might miss. They tailor preprocessing strategies for each script's unique visual complexities.

Terminology

Summary

The paper, LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4, addresses the enduring challenge of handwritten text recognition across diverse scripts, such as Arabic, English, and Persian.

Problem Statement

Accurately recognizing handwritten characters remains a complex challenge because each language and writing system introduces unique visual complexities and demands. Traditional Optical Character Recognition (OCR) approaches often rely on manually designed neural architectures, which require extensive expert intervention to iteratively adjust network layers, tune hyperparameters, or redesign model components. This process is described as being both time-consuming and prone to inconsistent outcomes.

Proposed Solution: The Automated Pipeline

The authors introduce a pipeline that is fully automatic and cross lingual, utilizing large language models (LLMs)—specifically GPT-5, GPT-4o, and Claude Sonnet 4—to act as independent designers for machine learning systems. This system requires no manual intervention, domain specific preprocessing, or human selection of models, resulting in a complete end to end automated system.

Methodology: The Closed Feedback Loop

The core of the proposed approach is an integrated and iterative loop comprising four key components: (1) data collection and augmentation, (2) LLM-driven neural architecture proposal, (3) automated model training and evaluation, and (4) continuous performance feedback for refining subsequent model architectures.

  1. Architecture Proposal: The pipeline begins with raw handwritten image data. Dataset-specific metadata is used to prompt the LLMs to propose candidate model architectures in structured JSON format. These specifications include a variety of components such as convolutional layers, pooling operations, normalization, dropout, and optional transformer-based modules, along with hyperparameters like the optimizer choice and learning rate.

  2. Model Construction: The system parses these JSON specifications and automatically convert[s] them into complete neural network models using Keras. This accommodates both conventional CNN-based structures and advanced hybrid architectures, including those incorporating VisionTransformer components.

  3. Evaluation: Each model is trained and evaluated without human intervention. Performance metrics—specifically training accuracy, validation accuracy, and test accuracy—are rigorously recorded.

  4. Refinement (The Loop): Following each trial, performance metrics are fed back into the same LLM that generated the architecture, creating a closed feedback loop. The LLM uses this information to refine its subsequent proposals based on previous outcomes, allowing it to adjust architectural elements such as layer types, depth, width, and hyperparameter settings in future iterations.

Scope and Application

The researchers applied this approach across three distinct scripts: Arabic (using the AHCD dataset), English (using EMNIST), and Persian (using the SADRI dataset). The entire process involved thirty independent trials for every language.

Key Contributions and Results

The key contributions of this work include a single, unified framework capable of recognizing multiple handwritten scripts, and the use of LLMs as standalone architecture generators that entirely remove the need for manual design.

The experimental results demonstrate that the pipeline consistently discovers efficient models:

  • Accuracy: The system achieved average scores above ninety three percent.

  • Efficiency: It maintained "inference speeds that meet the needs of real time applications. Notably, the system is able to automatically explore a wide range of neural architectures and adaptively select designs that fit the unique requirements of each script, without any explicit guidance from human experts."

The analysis showed that accuracy gains resulted from exploring depth and structural diversity rather than simply increasing parameter counts, validating the use of LLMs as independent designers for machine learning systems.

Improvements for AI systems

Based on a rigorous analysis of the provided research, I have identified several critical areas where this highly effective methodology can be advanced and expanded. The core strength—using LLMs as autonomous NAS agents in a closed-loop feedback system—is robust. However, its current scope is limited by specific deployment constraints and architectural generalizations.

Here are the improvements that can be made to enhance the system's capabilities, followed by what the resulting improved AI system can achieve.


The current evaluation focuses on mean latency and parameter count. This is sufficient for general comparison but insufficient for deployment on edge devices (e.g., mobile phones, embedded systems).

Improvement: Modify the LLM prompt structure to include Target Hardware Constraints. The input metadata will include specific hardware specs (e.g., Deploy on a TPU with 128MB memory or Requires sub-30ms latency). The LLM will then be instructed to optimize for these constraints, potentially choosing shallower, more efficient architectures over complex ones that were previously considered optimal purely based on raw accuracy.

The current system uses specialized data loaders for each script (EMNIST, SADRI, AHCD). This is manually defined by the a priori knowledge of the an AI researcher.

Improvement: Implement a LLM-Driven Adaptive Preprocessing Layer. Before the LLM generates an architecture, it will be prompted with detailed visual characteristics of new scripts (e.g., Devanagari script uses complex stacked characters). The LLM will then autonomously suggest the optimal preprocessing pipeline (e.g., specific normalization parameters, unique block-based segmentation) tailored to that script, making the system truly script agnostic and scalable to new languages.

LLMs are generative agents; they can generate syntactically valid JSON that translates into a mathematically flawed or computationally inefficient network topology (e.g, an infinite loop of layers, or incompatible activation functions).

Improvement: Introduce a Static Architecture Validator. This is a lightweight, deterministic module placed between the LLM output and the Keras construction phase. This validator will parse the JSON for logical consistency (e.g ensuring layer dimensions are compatible with subsequent pooling operations) and flag any structurally unsound or redundant configurations before they are ever trained, preventing wasted computational resources on hallucinated designs.

The current system achieves character-level recognition (classification of individual tokens). This is insufficient for real-world document processing, which requires recognizing sequences and understanding contextual flow.

Improvement: Expand the LLM's generative scope to include Hybrid Transformer/CNN Architectures. Instead of merely generating a stack of Conv2D/Dense layers, the LLM will be prompted to design architectures that incorporate self-attention mechanisms (Transformer blocks) after initial feature extraction. This allows the system to move from classifying individual characters into predicting sequences, enabling it to read full words and lines.

The integration of these improvements transforms the current model discovery pipeline into a truly autonomous, robust, and deployable AI solution:

  1. Achieve True Zero-Shot Multilingual Deployment: The system can be deployed to recognize any new handwriting script (e.g., Hindi or Thai) with minimal human intervention, as it autonomously determines the required preprocessing steps and generates a suitable architecture based on visual input characteristics, without needing a pre-trained dataset for that specific script.

  2. Guarantee Optimized Real-Time Performance: By incorporating hardware constraints into the search loop, the system guarantees that the discovered models are not only accurate but also optimized for specific deployment targets (e.g., "This model is guaranteed to run in <30ms on a mobile GPU"), making it suitable for real-world applications where latency is a hard constraint.

  3. Eliminate Computational Waste: The integration of the Static Architecture Validator ensures that every trial run is productive, as logically flawed or redundant architectures are discarded before training, maximizing the efficiency of the closed-loop search process.

  4. ** Deliver Contextually Aware Recognition:** By evolving from character classification to sequence prediction via hybrid LLM-generated architectures, the system will not only identify individual characters but also maintain and predict the correct ordering of words in a line, enabling it to handle complex handwriting and prepare data for downstream tasks like Optical Character Recognition (OCR) conversion into searchable text.

Sources

Related papers