On the arXiv today, with Purvam Jain, Preethi Jyothi and Vihari Piratla. https://arxiv.org/pdf/2605.30788 Modern LLMs have become increasingly good at a variety of languages. How does one detect cross lingual gaps in their abilities? We propose a set of synthetic puzzles that can be generated using base templates and then scaled up dynamically. This avoid translation errors; the tasks are clearly quantifiable; are comparable across languages; and as models improve it is trivial to scale up the test. The flip side is that these tests can’t detect more subtle issues like lack of awareness of nuance, or differences in creative-writing skill. Nevertheless, empirical experiments across many models, and seven languages — English, Hindi, Arabic, Chinese, Japanese, Tamil, Telugu — show that this benchmark can successfully detect cross lingual gaps. The complexity-accuracy graphs that we studied with Praneeth from a theoretical perspective in https://arxiv.org/abs/2601.14175 turned out to be useful. By studying the entire complexity-accuracy curve, we sidestep issues that arise if one tests the model at only a single complexity level.
