LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
summary
The gist
The paper, "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4," addresses the enduring challenge of handwritten text
In short
The episode discusses a paper titled "LLM-Driven AutoML for Cross-Lingual Handwritten OCR," which uses GPT-5, GPT-4o, and Claude Sonnet 4 to automatically design and refine neural network architectures for recognizing handwritten text in English, Persian, and Arabic. The hosts conclude that this closed-loop system removes the need for human intervention in model selection and tuning.
Key concepts
- LLM-Driven AutoML
- This involves using large language models like GPT to act as design agents that propose neural network structures for OCR tasks. These models generate specifications, such as JSON, which are then converted into working Keras models, automating the design process.
- Closed-Loop Neural Architecture Search
- This methodology means the system continuously tests and improves its own model designs based on performance results. It does not stop after one design but iteratively refines the architecture to find optimal structures for recognition.
- Cross-Lingual Handwritten OCR
- This refers to a system that can recognize handwritten text from different scripts, specifically English (using EMNIST), Persian (using SADRI), and Arabic (using AHCD). The paper focuses on creating a unified framework for these diverse languages.
- Autonomous Design Agents
- The use of multiple large language models—GPT-5, GPT-4o, and Claude Sonnet 4—as independent design agents allows the system to explore model architectures that human intuition might miss. They tailor preprocessing strategies for each script's unique visual complexities.
Terminology used across episodes
This episode discusses
- LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 · Paper Radio
- Qalam: A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
- GPT-NAS: Evolutionary Neural Architecture Search with the Generative Pre-Trained Model
- Can GPT-4 Perform Neural Architecture Search?
- GPT-4o System Card
- Handwritten Text Recognition: A Survey
The paper
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 · Read on arXiv
Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani
Iran University of Science and Technology Department of Computer Engineering, Tehran, Iran · Iran University of Science and Technology Department of Computer Engineering, Tehran, Iran
DOI: 10.1109/ICCKE68588.2025.11273810
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "LLM-Driven AutoML for Cross-Lingual Handwritten OCR".
Jane: The paper, "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we’re looking at the paper titled "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-five GPT-4o, and Claude Sonnet four" and the authors are Mobina Kashaniyan Amirhossein Ghassemi from Iran University of Science and Technology. It sounds like this paper is tackling the big problem of handwritten text recognition across different scripts using some very advanced language models.
Jane: That’s right, Tom, and it really highlights how these large language models can move beyond just generating text to becoming active participants in building complex machine learning systems for tasks like OCR. It suggests a new way of thinking about designing these kinds of tools.
Lu: I find the use of GPT-five GPT-4o, and Claude Sonnet four as independent design agents really interesting; it opens up possibilities for exploring architectures we might never even consider manually because human intuition has limits in that space.
Meng: From an engineering standpoint, the authors are focusing on creating a fully automatic pipeline that handles everything from data preparation to model training without needing constant human input for every step. That’s a big operational win if you’re trying to scale up these kinds of recognition systems.
Lalam: The title itself points to this "Closed-Loop Neural Architecture Search," which implies the system doesn't just build one thing and stop; it keeps testing and improving the design based on its own performance results. That kind of continuous self-improvement feels very promising for cultural preservation tasks, which is what this whole project is focused on.
Tom: Exactly, Jane; that closed loop means the system gets smarter with every trial, which moves it away from static designs we use today. This paper is showing how to automate that iterative refinement process using the power of these large language models. Where do we go from here?
The paper's summary: Jane: Now that we know the setup, let’s look at what the paper actually describes in terms of its core methodology. Essentially, they are taking raw handwritten images from different sources like English with EMNIST, Persian with SADRI, and Arabic with AHCD.
Tom: And what they do is feed those images into a pipeline where the AI models—GPT-five through Claude Sonnet four—are tasked with proposing the actual neural network structures for recognition. It’s not just guessing; it’s generating specifications in a structured format, like JSON, which then get converted into working Keras models.
Lu: What I find particularly compelling is how they tailor the preprocessing and encoding strategies specifically for each script to match its unique visual complexities, which addresses the initial problem statement about different languages demanding unique solutions.
Meng: They also mention using a lightweight data augmentation strategy, like random rotations and shifts, during training with Keras’s ImageDataGenerator to help the model generalize better without adding too much computational load during the search phase. That shows they are thinking about practical deployment constraints right from the start.
Lalam: This whole summary points to a system that is cross-lingual by design, meaning it’s not just a set of separate tools for English and Arabic, but one unified framework that handles them all within the same automated search process.
Tom: So, they are using these LLMs as the primary designers to propose architectures that fit the specific data they feed them, which is a major departure from how we typically design neural networks today. Jane, what does this mean for practical implementation?
The paper's improvements: Jane: The main improvement they highlight is completely removing the need for human intervention in model selection and hyperparameter tuning; instead, the system handles that entire iterative optimization cycle automatically. This eliminates a major bottleneck that used to slow down development significantly in traditional OCR work.
Tom: That’s a huge operational advantage, Jane; being able to scale up across many languages without needing a team of experts for every single model design is something we need to focus on for real-world adoption. Lu, what about the performance aspects?
Lu: The authors show that the system leverages optimal architectural patterns and hyperparameters specific to each targeted handwriting script, which means they aren't just randomly picking structures; they are using knowledge of how those scripts are visually organized.
Meng: They also find that this automated process maintains inference speeds that meet the needs of real-time applications, and they noted that the system explores a wide range of neural architectures adaptively based on what it learns during the search. That adaptation is key for efficient deployment.
Lalam: This capability to automatically select designs that fit the unique requirements of each script, without explicit human guidance, really speaks to how adaptable these AI systems can be when applied to diverse cultural data sets.
Tom: It seems like they've moved the focus from designing a single perfect model to building a scalable system capable of discovering many efficient models tailored for many different languages. This paper is showing a new path for AI development in this area.
Conclusion: Jane: So, to wrap up our discussion on "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-five GPT-4o, and Claude Sonnet four" the core implication is that these LLMs are acting as autonomous generators and evaluators rather than just passive tools.
Tom: Exactly, Jane; it's about a shift from relying on human intuition for design to having an autonomous system discover and refine the best architectures through a closed-loop process. This whole paper shows how effective this approach can be when applied across multiple languages simultaneously.
Lu: The potential for these models to discover novel architectures is something we will definitely see expanded upon in other areas, perhaps in analyzing complex visual patterns far beyond just text recognition.
Meng: For my team, the practical implication is that this gives us a framework for building systems that adapt to real-world handwriting complexity without being constrained by pre-defined designs or needing constant manual tuning cycles from human experts.
Lalam: We can use this to ensure that historical records aren't lost because they're too complex or expensive to digitize using conventional methods, making knowledge more accessible globally through this technology.
Tom: This entire journey from manual model building to seeing "LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-five GPT-4o, and Claude Sonnet four" is a significant step forward in how we do AI research.
Jane: It’s clear that the system is achieving both robust performance across Arabic, English, and Persian while maintaining good efficiency metrics which makes this a very compelling piece of research.
Lu: I think the potential for finding optimal patterns across different language structures suggests that the very definition of how we design AI systems is changing here.
Meng: It provides a clear pathway to rapid and adaptable deployment of OCR technology across many languages and domains, addressing real-world operational needs that were previously unmet by slower methods.
Lalam: We can use this to ensure that historical records aren't lost because they're too complex or expensive to digitize using automated means.
Tom: Thanks everyone for joining me on the show; we’ll be right back with the next paper from arXiv!
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization