ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

arXiv:2606.20179 · cs.CL · Submitted 2026-06-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion".

Jane: Grapheme-to-phoneme (G2P) conversion for Modern Hebrew is needed for applications like text-to-speech (TTS), but is challenging due to the language’s abjad writing system, which leaves vowels largely unwritten,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we’re looking at this paper now called "ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion." Essentially, the core problem they're tackling is how to convert written Hebrew into actual spoken sounds, which is really tough because the writing system leaves out most of the vowels.

Jane: Exactly, and this paper proposes a new way around that challenge by using audio supervision to guide the process.

Lu: It’s interesting because they are addressing two major weaknesses in current methods: they're moving beyond just vowel diacritics to include things like lexical stress, which standard approaches miss entirely.

Meng: That makes sense from an engineering standpoint; knowing where the stress falls is crucial for natural-sounding text-to-speech.

Lalam: I think this approach could really open up how we process and understand spoken Hebrew on a much deeper level for cultural applications.

Tom: Right, and what they claim is that ReNikud uses a two-stage pipeline to get around those limitations, starting with generating pseudo-labels from massive amounts of unlabeled audio.

Jane: They build this pipeline using two different Automatic Speech Recognition systems running in parallel—one standard and one custom—to create character-aligned IPA transcripts from the audio.

Lu: That step is clever because it uses both orthographic and phonetic transcriptions to align things, which is a big deal when dealing with the abjad structure.

Meng: So, they are essentially using the ASR outputs as a proxy for what the correct pronunciation should look like, even though those labels aren't perfect.

Lalam: It’s about leveraging scale; by applying this to thousands of hours of Hebrew audio, they can gather a lot of data that would otherwise be impossible to collect manually.

Tom: And then they move into their main architecture, which is called the pseudo-vocalization architecture, where each character predicts its phonetic components independently.

Jane: That part sounds intricate because instead of predicting a whole sequence at once, every single letter has to make three separate predictions: one for consonant, one for vowel, and one for stress.

Lu: The use of three parallel classification heads—one for consonants selecting from twenty-five symbols or null, another for vowels choosing between five vowels or null plus a special token like /aχ/, and a third specifically for lexical stress—gives the model fine-grained control.

Meng: From an implementation view, having those independent heads allows the model to focus its attention on different phonetic features simultaneously at the character level.

Lalam: This level of detail is what I think will allow us to build text-to-speech systems that capture spoken Hebrew nuances much more accurately than before.

Tom: So, moving into their evaluation, they tested this framework on existing benchmarks and introduced a new one called MILIM to see how well it handles tricky situations.

Jane: The authors found that ReNikud performs better than the baseline model across all categories on MILIM, specifically showing "the largest gains on categories reflecting spoken Hebrew norms."

Paper summary: Lu: That result is significant because it means their method actually captures the specific way people speak Hebrew in real-world contexts, like slang or loanwords.

Meng: I'm interested in how much better it performs when dealing with those colloquialisms compared to the standard sequence-to-sequence models they compared against.

Lalam: It shows that audio supervision isn't just for formal speech; it’s effective at modeling the spoken reality of the language, which is a huge step forward for language understanding.

Tom: Furthermore, they showed that this character-level encoder can actually be adapted to help with traditional Hebrew text diacritization, achieving competitive accuracy using "substantially less training data than standard approaches."

Jane: That comparison against standard methods highlights the efficiency of their approach when you have limited resources for training.

Lu: The ablation studies they conducted were very telling; they confirmed that the character-level alignment acts as a necessary guide, because without it, sequence-to-sequence models frequently make mistakes on vowels and stress placement.

Meng: That confirms that forcing the model to respect the structure of the abjad is what prevents it from just guessing phonetic features randomly.

Lalam: This suggests that for applications where data is scarce, leveraging structural knowledge like character alignment provides a much stronger inductive bias for learning correct pronunciation patterns.

Tom: Looking ahead at what this means for the future, they’ve shown that audio-derived labels outperform text-derived labels across all categories on MILIM.

Jane: That points toward a future where we can train these systems using readily available speech data rather than relying on painstakingly transcribed text, which is a big practical advantage.

Lu: The transferability of the framework to other languages with similar challenges, like Arabic, is also something they mentioned as promising for broader impact.

Meng: Practically speaking, if we can train models this way efficiently using audio labels, it lowers the barrier significantly for developing high-quality speech technology in many less-resourced languages.

Lalam: For culture and accessibility, this means that spoken Hebrew content will be more accurately converted to text or synthesized back into speech with much higher fidelity than what we could achieve otherwise.

Tom: So, to wrap up these points on the ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion paper, we see a method that uses parallel ASR for pseudo-labeling and a constrained classification architecture to predict character triplets like consonant, vowel, and stress.

Jane: It tackles the ambiguity of unwritten vowels by framing the problem as per-character classification with three specialized heads.

Lu: The success on MILIM demonstrates its ability to handle spoken Hebrew norms better than previous sequence-to-sequence models when dealing with slang and loanwords.

Meng: From an engineering standpoint, it shows a way to achieve good results without requiring massive amounts of perfectly annotated training data initially.

Lalam: Ultimately, this work provides a powerful tool for improving text-to-speech and other speech technologies for Hebrew by modeling how the language is actually spoken, which will make those applications much more realistic.

Conclusion: Tom: So, to wrap up what we’ve heard about ReNikud, we’re talking about this paper that tackles the tough job of converting written Hebrew into actual spoken sounds using audio information and a smart new architecture.

Jane: Exactly, Tom; it’s all about taking those tricky vowels and stresses in Modern Hebrew and mapping them correctly by listening to how people actually speak it.

Lu: The authors really managed to build a system that uses parallel speech recognition models to create these pseudo-labels, which is a clever way to get data where it's hard to find.

Meng: I’m curious about the practical side here; if this works well for Hebrew, does that mean we can actually make text-to-speech for this language sound much more natural for everyday use?

Lalam: From my view, the potential impact is huge because it unlocks a lot of cultural content previously stuck in written form, allowing us to bring spoken Hebrew closer to high-quality digital experiences.

Tom: Right, and the authors introduced this ReNikud framework specifically because existing methods struggle with that ambiguity in the first place.

Jane: It really does; they show how framing the task as a constrained classification problem for each character helps solve those vowel and stress prediction issues we’ve seen before.

Lu: Their finding that audio-derived labels are better than text-derived labels across challenging categories like slang is something I think opens up some really interesting avenues for future research in low-resource language processing.

Meng: It seems like the main implication is a more robust and efficient way to handle these complex phonetic details without needing massive, perfectly transcribed datasets upfront.

Lalam: And for culture, this means that historical texts or even new digital media can be brought back to life with a level of phonetic accuracy that was just out of reach before.

Tom: It really sounds like the title itself, ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion, sums up exactly what this paper is achieving.

Jane: It’s a very precise description of the method—using audio supervision to figure out those grapheme-to-phoneme conversions for Modern Hebrew.

Lu: I think the real excitement is seeing how they adapted the character-level encoder for traditional text diacritization, which shows this isn't just a niche solution but a more generalizable technique.

Meng: It’s interesting that their ablation studies confirmed the importance of that character-level alignment; it shows you can't just throw a big sequence model at it and expect good results.

Lalam: So, we’re looking at a tool that bridges the gap between how Hebrew is written and how it actually sounds in conversation, which could fundamentally change how we interact with digital versions of the language.

Tom: And that leads us perfectly into what happens next—we’ll be digging deeper into those specific benchmarks they used to validate these results.

Reichman University · Carnegie Mellon University

cs.CL

Submitted: 2026-06-18

Updated: 2026-09-30

Journal ref: SLT IEEE 2026

Code: https://github.com/maxmelichov/BlueTTS

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: Grapheme-to-phoneme (G2P) conversion for Modern Hebrew is needed for applications like text-to-speech (TTS), but is challenging due to the language’s abjad writing system, which leaves vowels

Key concepts

Audio Pseudo-Labeling
This pipeline uses two parallel Automatic Speech Recognition (ASR) systems—one for standard text and one custom model for IPA transcription. By aligning these outputs based on Hebrew orthography, it creates 'pseudo-labels' that link written characters directly to phonetic transcriptions from unlabeled audio.
Pseudo-Vocalization Architecture
Instead of predicting complex diacritics, this core model treats G2P as a classification problem for every single letter. Each letter independently predicts one phonetic triplet: a consonant, a vowel, and a stress marker. This character-level prediction is guided by three separate classification heads (consonant, vowel, stress).
Character-Level Transformer Encoder
This component processes the input text at the individual character level. It uses parallel classification heads—one for consonants (25 options), one for vowels (5 options plus null/special tokens), and one for lexical stress (binary)—to generate the phonetic predictions simultaneously.

Terminology

Summary

Grapheme-to-phoneme (G2P) conversion for Modern Hebrew is needed for applications like text-to-speech (TTS), but is challenging due to the language’s abjad writing system, which leaves vowels largely unwritten, creating substantial ambiguity.

The gist

ReNikud proposes a framework that overcomes limitations in Hebrew G2P by using weak audio supervision via an ASR pseudo-labeling pipeline and a pseudo-vocalization architecture that enforces character-level alignment.

How it works: Audio Pseudo-Labeling (Section II-A)

To learn pronunciation from unlabeled audio at scale, the method constructs a pipeline that extracts character-aligned IPA annotations using two parallel ASR systems. One system extracts Hebrew orthographic transcripts with a standard pretrained Hebrew ASR model, while the other custom ASR model is trained to output IPA when applied to Hebrew audio. By applying both systems to a large-scale, unlabeled Hebrew audio corpus, the pipeline yields parallel Hebrew text and IPA transcriptions that serve as pseudo-labels. To find character-level correspondences between orthographic and IPA transcripts, a string alignment process based on the Hebrew orthography’s abjad structure is performed. This process uses a simple finite state transducer (FST) to handle complexities such as digraphs (e.g., using an apostrophe for non-native phonemes), word-final /χ/ sounds with preceding vowels, and silent letters, mapping graphemes to phonetic triplets encoding consonant, vowel, stress.

How it works: Pseudo-Vocalization Architecture (Section II-B)

The core of the method is a pseudo-vocalization architecture that predicts IPA phonemes at each character position. This approach frames the G2P task as a constrained, per-character classification problem. Instead of predicting traditional diacritics, it directly predicts character-aligned IPA phonemes by having every Hebrew letter independently predict exactly one phonetic triplet (consonant, vowel, stress). The model utilizes a character-level transformer encoder with three parallel, independent classification heads: a Consonant Head (selecting from 25 IPA consonants or null), a Vowel Head (selecting from 5 vowels, null, or the special /aχ/ token), and a Stress Head (binary classifier for lexical stress). At inference time, realizations are predicted by taking the argmax of logits for each head, with constrained decoding applied to enforce hard constraints on Hebrew letters and phonetic realizations.

Evaluation and Benchmarking (Section III)

The method is evaluated on existing Hebrew G2P benchmarks and a new targeted benchmark called MILIM. The MILIM5 benchmark is introduced to assess models' ability to perform complex phonetic disambiguation in challenging Hebrew contexts, such as slang, foreign loanwords, acronyms, and colloquial speech. ReNikud outperforms the baseline model across all categories on MILIM, showing the largest gains on categories reflecting spoken Hebrew norms. For example, it correctly predicts colloquial forms like החפשמו → /vemiʃpaχˈa/ where other models default to the prescriptive form. Furthermore, the framework demonstrates broader utility by adapting the character-level encoder to traditional Hebrew text diacritization, achieving competitive accuracy with substantially less training data than standard approaches.

Ablations and Transfer (Section F)

Ablation studies confirm the necessity of key components. Comparing architectures, ReNikud’s constrained version outperforms both a seq2seq baseline and a CTC network. The paper demonstrates that without character-level alignment as an inductive bias, sequence-to-sequence models frequently mispredict vowels, stress placement, and other underspecified phonetic features. Data source comparisons show that audio-derived labels outperform text-derived labels across all MILIM categories, with the largest gains on colloquial and slang items where spoken and written norms diverge most. The transfer to diacritization shows that the audio-supervised pretraining provides an inductive bias that guides correct vocalization, allowing the model to converge significantly faster at extremely small data sizes compared to baselines.

Limitations and Future Work (Section V)

A primary limitation is the reliance on ASR-generated pseudo-labels, as inherent transcription errors from the ASR model naturally propagate into the G2P training data. Additionally, because the training corpus is skewed toward formal discourse (Knesset Vox), further work could incorporate more diverse conversational audio. Finally, the framework shows promise for additional languages sharing similar challenges, such as Arabic. The paper also notes that its alignment mechanism tends to discard highly informal spoken variants that deviate significantly from the written orthography.

Index Terms

Grapheme-to-Phoneme, Text-to-Speech, Modern Hebrew, Weakly Supervised Learning, Automatic Speech Recognition, Lexical Stress.

--- Page 6 ---

Table III: Detailed G2P evaluation on MILIM (WER / CER, in %).

Improvements for AI systems

Here are specific improvements that can be made to existing AI systems, based on the methodology and findings of the ReNikud paper:

  1. The core G2P model architecture should be reframed from a standard sequence-to-sequence (seq2seq) prediction task into a per-character classification problem, termed Pseudo-Vocalization.

  2. Instead of predicting traditional diacritics, the system must directly predict character-aligned IPA phonemes as a phonetic triplet: (Consonant, Stress, Vowel) for every single Hebrew grapheme.

  3. The model should utilize an encoder (like DictaBERT) paired with three independent classification heads—one for consonants, one for vowels/special tokens (/aχ/), and one binary classifier for lexical stress—to predict these triplets simultaneously at each character position.

  4. The system must enforce a strong inductive bias by using a Finite State Transducer (FST) during the alignment phase to map unvocalized Hebrew characters to their corresponding phonetic triplets, explicitly handling orthographic complexities like loanword digraphs (geresh), word-final stress patterns (patah gnuva), and silent letters as specific classes.

  5. To leverage abundant unlabeled audio data, implement a weak supervision pipeline using two parallel ASR systems: one trained on standard Hebrew orthography and one custom ASR trained specifically to output IPA when processing Hebrew audio. These parallel outputs are then aligned via the FST to generate pseudo-labels for training the G2P model.

  6. Inference, instead of relying on unconstrained decoding (which leads to poor performance), apply constrained decoding: only allow predictions over the set of possible consonantal realizations for a given letter and enforce exactly one lexical stress prediction at word level.

These improvements will enable an AI system that can perform the following specific functions:

  1. Perform high-accuracy, character-level Grapheme-to-Phoneme (G2P) conversion for Modern Hebrew text, moving beyond the limitations of vowel diacritics alone.

  2. Generate accurate International Phonetic Alphabet (IPA) transcriptions directly from unvocalized Hebrew script.

  3. Identify and correctly resolve complex phonetic ambiguities in spoken Hebrew, including those arising from homographs, slang, loanwords (via geresh), and colloquial pronunciation that deviates from formal written norms.

  4. Enable robust Text-to-Speech (TTS) systems for Hebrew by providing the precise phonetic features required for natural prosody generation, as the model will explicitly predict lexical stress placement.

  5. Transfer learned representations from audio supervision to high-accuracy diacritization tasks, allowing diacritic prediction on severely data-scarce text corpora with competitive performance.

Sources

Related papers