Overcoming Vocabulary Constraints with Pixel-level Fallback
AuthorsJonas F. Lotz†**, Hendra Setiawan, Stephan Peitz, Yova Kementchedjhieva‡
Overcoming Vocabulary Constraints with Pixel-level Fallback
AuthorsJonas F. Lotz†**, Hendra Setiawan, Stephan Peitz, Yova Kementchedjhieva‡
Subword tokenization requires balancing computational efficiency and vocabulary coverage, which often leads to suboptimal performance on languages and scripts not prioritized during training. We propose to augment pretrained language models with a vocabulary-free encoder that generates input embeddings from text rendered as pixels. Through experiments on English-centric language models, we demonstrate that our approach substantially improves machine translation performance and facilitates effective cross-lingual transfer, outperforming tokenizer-based methods. Furthermore, we find that pixel-based representations outperform byte-level approaches and standard vocabulary expansion. Our approach enhances the multilingual capabilities of monolingual language models without extensive retraining and reduces decoding latency via input compression.
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
August 20, 2026research area Speech and Natural Language Processing
Cross-lingual knowledge transfer is critical for building high-performing multilingual language models for languages with insufficient training data. When target language data is scarce, the knowledge required for many downstream tasks involving scientific reasoning, commonsense inference, and world knowledge must be acquired primarily from the high-resource language, making effective knowledge transfer essential. Existing methods for improving…
Cut Your Losses in Large-Vocabulary Language Models
February 7, 2025research area Methods and Algorithmsconference ICLR
As language models grow ever larger, so do their vocabularies. This has shifted the memory footprint of LLMs during training disproportionately to one single layer: the cross-entropy in the loss computation. Cross-entropy builds up a logit matrix with entries for each pair of input tokens and vocabulary items and, for small models, consumes an order of magnitude more memory than the rest of the LLM combined. We propose Cut Cross-Entropy (CCE), a…