Here's an uncomfortable question: if you recorded every last word of a disappearing language and stored it perfectly forever, would you have saved it?
No. You'd have a very well-organized tombstone.
I've been thinking about this because of how AI language models get trained. There's a real fear going around — that AI is flattening culture, favoring big languages like English or Mandarin, and quietly erasing smaller ones. That fear is legitimate. But most of the proposed fixes miss what's actually happening.
The AI isn't the disease. It's a magnifying glass.
Big languages dominate training data because big languages already dominate the internet, publishing, and education systems. That imbalance existed long before any chatbot. AI didn't invent it — it just processes what's already out there, faster and at greater scale. So when a model sounds a little generic, a little "averaged out," it's reflecting the world's existing imbalance back at us. Blaming the mirror won't fix the room.
Here's the assumption almost everyone makes without noticing: more data equals more preservation.
It feels true. It isn't.
A language survives because people use it — with their kids, at the market, in arguments, in jokes. It survives through daily, living use between real people. Storage — recordings, archives, searchable text — is just a byproduct of that use, not the thing that keeps it alive.
I've seen well-funded projects build gorgeous digital archives of a language that, fifteen years later, no child speaks at home. The archive is beautiful. The language is still dying. Because nobody fixed the actual reason people stopped passing it down — usually money, migration, or a lack of any everyday reason to keep speaking it.
So what actually matters?
Three things, and they're simpler than most strategies admit:
The people who carry a culture forward are irreplaceable, and they're running out of time. Once the last person who can teach a specific song, ritual, or skill passes away without teaching someone else, that's it. No amount of computing power brings it back. This is the single most urgent fact in the entire conversation, and it gets the least attention.
Communities keep a language alive when there's a real reason to use it — not because someone archived it. Paid work, local status, entertainment, commerce. If speaking a language helps someone provide for their family or feel proud in their community, it survives. If it doesn't, no archive changes that.
Badly organized cultural material can't be honored properly by anyone, ever — AI or otherwise. Even if AI weren't in the picture, undocumented, unlabeled cultural knowledge is fragile. Getting this right is basic housekeeping, not a technology solution.
What I'd actually do, starting this week, if this mattered to me:
Skip the debate about global AI training data — you have almost no leverage there, and campaigning for it takes years. Instead, find one living person who still carries a specific piece of cultural knowledge — a craft, a story, a way of speaking — and one person willing to learn it from them. Record the actual teaching happening, not just an interview about it. Write down why it matters and when it's used. And before any of it goes online anywhere, decide who owns it and whether it can ever be used to train an AI model at all.
That single recording, done with intention, outlasts a thousand well-meaning archiving grants.
The real fix isn't a better dataset. It's making sure someone still wants to teach the next person before there's no one left who can.
Curious what others working in heritage preservation, linguistics, or AI ethics think — is data governance for cultural material something communities are actually equipped to negotiate right now, or is that its own uphill battle?
