Modern language models are extraordinarily capable of generating text - fluent, endless, eager to say everything. So I asked the opposite question:
What happens when you take an LLM's entire vocabulary away?
Strip a model down to a tiny, fixed vocabulary and ask it to carry a whole inner life through the gaps. Can a character still be expressive? Can it still reach you with only a handful of words to her name? Can Lily tell the world her story when she's barely allowed to speak?
So Lily has no words. She has 28 tones - I, you, music, family, remember, afraid, before, nothing, friend, lonely ... and nothing else. When you ask "do you like music?", she might answer ◈ ◉ ♪ which means I · feel · music and it's on you to hear that as "yes, it's everything to me." The lexicon panel is your notebook: you write down what you think each glyph means, and a model quietly grades how close you are. The game is the slow joy of a language coming into focus.
GAMEPLAY
1 · You ask. She answers in tones. Type a question into the interrogation box - "Why are you afraid?" Lily replies with a short chord of glyphs that plays back as a melody on a music box. No subtitles. The first time, it's just sound and symbols.
2 · You guess what each glyph means. Open the lexicon panel which is your notebook for this case. Click a glyph you keep hearing and write your best guess: "this one feels like… remember?" A model silently grades it and warms the glyph's colour from red → amber → teal → green as you close in. You're never told the answer; you feel your way to it.
3 · The language comes into focus and so does the story. As more glyphs resolve, her answers stop being noise and start being sentences. ◈ ◉ ♪ stops being three shapes and becomes I · feel · music. Evidence unlocks. New tones unlock. The mystery of what happened to Lily tightens with every word you decode, until the last act, when she speaks once, aloud, in a voice you've spent the whole game earning the right to hear.
ARCHITECHTURE
Lily's vocabulary is essentially a small langauge model which is Qwen2.5-7B. The model is only allowed to answer from fixed set of vocabularies in response. Furthermore, due to the nature of the game it was imperative that I add a judgement layer to filter out the nonsense inputs that might come from users. So Nemotron-mini-4B acts as the judge to filter out and reject unwanted inputs. The same model aslo acts as the judge of the user's guess for lily's tones.
The app was built entirely on Gradio and was later hosted on Hugging Face Space. The models were hosted on Modal.
AWARD
I built Forgotten Lily as part of Hugging Face's Build Small Hackathon where you were only allowed to use models below 32B parameters in accumulation. While I did not expect to win this hackathon by no means, a month later, to my surprise, they awarded me an award for the best demo!
Awarded $1000 USD
