language revitalization · ai strategy lead
Wikitongues
AI that speaks endangered languages - and the first community-owned benchmark for it.
Every number on this desk carries its date. The scoreboard below asks the project database for fresh ones while you read.
I'm the AI strategy lead of Wikitongues, a nonprofit working to sustain and document every language in the world. The initiative I run builds AI that speaks endangered languages authentically, and the first community-owned benchmark for how well frontier models actually speak them.
the benchmark is the lever
About 7,000 languages are spoken today - roughly half endangered, under 5% represented in AI. Off-the-shelf models answer confidently and get them wrong. A community-owned leaderboard, measuring how well ChatGPT, Gemini, and Claude actually speak a given language, is the lever that pushes the labs to do better.
the igala pilot
The first pilot is Igala (Yoruboid, around two million speakers, Kogi State, Nigeria), led in-community by Agnes Abah, community lead of the Igala Wikimedians. Today's frontier models fail Igala badly, often confusing it with neighbouring languages.
funding and the ghana launch
$25,000 raised for the first three months, with collaboration and support from Google Research (Impact Lab), DAIR, and Georgia Tech, alongside academic advisors at NYU, JHU, and Sydney. The roadmap: a public launch and the first-ever Igala leaderboard at the Wikimedia Foundation conference in Ghana, October 4 to 5, 2026.
Both views show the public site, not the annotation platform where the speakers work. Platform screenshots are not on this desk yet.
the language
Igala is a tonal language of around two million people in Kogi State, Nigeria. It has seven vowels, two of them written with a dot underneath (ẹ and ọ). A changed letter is a different word, not a typo.
what the models know
Igala is not in any of the major multilingual AI training or evaluation datasets. The models are not bad at Igala. They have almost never seen it.
what they do instead
Ask most AI models a question in Igala and they answer in Yoruba or English instead, with invented words in between. A model that does not know Igala does not say so. It guesses a nearby language. One speaker, shown a wrong answer for the word "morning", said: it's not an Igala word, maybe it's coming from Yoruba.
how bad, counted
In the first 781 blind comparisons, all of them on GPT-4o and GPT-4.1 class models, native speakers rejected both answers 775 times. Five comparisons picked a side and one was called a tie.
counted 2026-08-09Two questions from the frozen exam, answered by the same models twice: once asked plainly, and once with the community's own Igala placed in front of them. How that placing works is the rest of this desk. The community's answers are Omi and Oji.
before · a plain model
- the igala word for water
- ámẹ́ · amá · màíí
- the igala word for husband
- úchu · Ǹdá · Àgbá
after · with the knowledge packed in
- the igala word for water
- Ómi
- the igala word for husband
- Oji
The plain answers are Yoruba-flavoured guesses. With the packed bundle every retrieval version wrote Ómi, the fine-tuned model wrote Omi, and husband followed the same pattern. On a dialect question one plain model answered in Igbo.
recorded 2026-08-09, on the frozen examFour layers, read top to bottom. The community produces the knowledge. The knowledge is packed around each question. A model answers. And every answer goes back to the community for judgment, which becomes new knowledge.
The small red mark between layers two and three is the leak guard. It checks every retrieved piece, so no exam question is ever handed its own answer. That check is the reason the scores below can be believed.
Every judgment and every correction re-enters the knowledge, the grammar, and the next round of models. The community is not labeling data for a system. The community is the system.
When someone asks a question, the system packs a bundle around it, in this order, and sends the whole bundle to the model. Today's version is v4.1.
On exam questions, every piece passes the leak guard first: if a piece contains that question's own community answer, it is dropped and the drop is recorded. Otherwise the test would hand the model its answer key.
1 · the rules none how to use everything below v2
A numbered procedure telling the model how to use everything below: the dictionary for word forms, the examples for sentence shape, and, since v4, one rule above all others: translate the meaning of the whole sentence, never word by word.
2 · real answers by speakers leak guard what a good answer sounds like v1
Question-and-answer pairs written by Igala speakers, shown as example exchanges. The model sees what a good answer looks and sounds like: short, in Igala, spelled the community's way.
3 · corrections speakers made leak guard the mistake and the fix, side by side v4
Since v4: a few cases where a model wrote something, a speaker fixed it, and the speaker said why. The model sees the mistake and the fix side by side, which teaches more than the fix alone.
4 · example sentences leak guard how Igala sentences are built v2
Igala-English sentence pairs that show how Igala sentences are built. Served only when the question asks the model to build something, a sentence or a story or a greeting, because word-lookup questions were measurably hurt by them.
5 · dictionary lines leak guard the exact attested form of each word v2
One line per meaningful word of the question, with the exact attested Igala form. Placed right above the question because in Igala spelling is meaning: a changed letter is a different word, not a typo.
6 · the question, then one closing rule none answer in Igala only, nothing else v4.1 repair
The person's actual question, unchanged. Under it, a single line restating the output rule: answer in Igala only, nothing else. Then, on v4.1 only, a check of the finished answer. If it uses letters Igala does not have, or is crowded with tone marks the community would not write, the model is asked once to rewrite it. So a v4.1 answer can be the model's second attempt, and its exam score is scored that way.
None of these models learns Igala the way a person does. Each version changes what real Igala the model gets to see at the moment it answers, and how it is told to use it. Each fix exposed the next problem.
v0 - plain models baseline Ask a top model, nothing added on the board
what it fixed
Nothing yet. This is the baseline.
what it did not
Asked for Igala, the models answer in Yoruba or English, with invented words in between.
v1 - show it real answers aug 9, 2026 Community answers pasted into the prompt retired
what it fixed
Real Igala words start appearing.
what it did not
Words without sentence structure. A community reviewer put it plainly: the first sentence is saying three different things.
v2 - give it a method aug 12, 2026 A dictionary, example sentences, and a step-by-step procedure retired
what it fixed
Correct spellings, and sentence shapes copied from real ones.
what it did not
Still copying, not speaking.
v3 - teach it the grammar aug 14, 2026 Everything in v2, plus grammar rules read out of the evidence in blind judging
what it fixed
Pronouns, negation and word order arrive as rules. Speakers prefer this version to the plain model when judging blind.
what it did not
The same rules made Claude worse, not better. And on the exam, v3 is not measurably ahead of the plain model.
v4 - translate the meaning aug 29, 2026 The instructions rewritten around one rule from the community's review unjudged
what it fixed
Whole-sentence meaning instead of word by word. Speakers' own corrections are now packed into the prompt too.
what it did not
Long, cultural, open-ended questions still collapse. No speaker has judged this version yet.
v4.1 - the rules the failures taught us aug 31, 2026 v4 plus rules mined from every judged failure, and a repair round serving today
what it fixed
Undoes the damage v3 did to Claude. Beats v3 on the exam in a paired test, the only step between versions that does.
what it did not
The gain over v4 is mostly fewer tone marks, which the community rarely writes. No speaker has judged this version yet.
the questions
Every model takes the same exam: 43 frozen questions that no model is served the community's answer to, and that no speaker is ever asked to judge. Each answer is compared with what Igala speakers wrote for the same question. 16 of the 43 were once handed one of their own community answers in the material served to a model, so every published score uses only the 27 where that never happened.
counted 2026-09-03the score
Underneath is chrF, the standard overlap score used to grade machine translation: 0 to 100 for how much an answer's letters and letter-pairs overlap with the community's answers. It is measured on the answer itself, so an English preamble cannot inflate it. The score is then rescaled so that the agreement between two native speakers reads exactly 100. A score of 85 means: this model's answers are 85% as close to the community's writing as one speaker's answers are to another's.
why 100 means the speakers
Two Igala speakers answering the same question rarely write the identical string. On the leak-free questions, one answer per speaker, their agreement is 39.5 chrF. That number is the 100 line. Say a question asks for one word, and two speakers wrote the same five letters with one accent mark different. Their overlap is high but not perfect, and that overlap is what the 100 line is anchored to. A model that shares four of those five letters lands near the line. A model that answers in English shares almost nothing and lands near zero.
ceiling computed 2026-09-03what a bar past 100 means
A bar past the 100 line does not mean the model beat the speakers. The model is scored against every community answer for a question, while each speaker is scored against the other speakers only, and that gives models a built-in advantage that grows with the number of answers per question. Scored like-for-like, the best system sits at speaker level.
what the exam cannot see
The score measures resemblance to how the community writes. Only native judgment measures fluency. Most of the 43 questions ask for a word or a short phrase, so the exam cannot register how a model handles a greeting or a long answer.
how a judgment works
A speaker sees two answers to one question without knowing which system wrote which. She picks the better one, calls it a tie, or rejects both as inadequate. She tags what went wrong and may correct the winner.
the one result
Since August 20, 2026, every blind judgment in the comparison pool compares Gemini 3.1 Pro with the v3 package against the same Gemini with nothing added. Speakers prefer the package: 54 wins to 14, with 21 ties and 104 pairs where both answers were rejected. Counted once per question, it is 25 to 7 over 55 questions. Every one of the six annotators leans the same way.
193 comparisons, counted 2026-09-01both inadequate
Across all 1,295 blind comparisons recorded so far, speakers rejected both answers 1,159 times. On the current pairing the rate is about 48%, 119 of 247. In the first 781 comparisons it was 99%, and those were judged on GPT-4o and GPT-4.1 class models. The fall came with a change of models, not from the method.
counted 2026-09-03an ai judge was tried
A model was asked to grade the pairs instead of a speaker. It agreed with the humans 90% of the time, but only because both mostly said no. Once that is accounted for it was at chance. The judge does not know Igala either, so it is kept as an alarm, not a ranker.
781 comparisons, measured 2026-08-09what no one has judged
No speaker has judged v4 or v4.1, the versions that top the exam. That test comes before any further prompt work.
Community Agreement Score per arm. An arm is one model with one version of the package, or with nothing added. A higher number means closer to how the community writes. 100 is the line where two native speakers agree with each other.
snapshot of the live board, computed 2026-09-03 21:37 utcGemini 3.1 Pro + Igala RAG v4.1 120.1 the v4.1 package · 93.2 to 148.9 27 of 43
Raw leak-free chrF 47.5. Rank 1 of 26 arms.
Gemini 3.1 Pro + Igala RAG v4 102.1 the v4 package · 73.4 to 133.4 27 of 43
Raw leak-free chrF 40.3. Rank 2 of 26 arms.
Gemini 3.1 Pro + Igala RAG v3 99.2 the v3 package · 74.3 to 125.9 27 of 43
Raw leak-free chrF 39.2. Rank 3 of 26 arms.
Gemini 3.1 Pro 94.2 nothing added · 68.5 to 122.1 27 of 43
Raw leak-free chrF 37.2. Rank 4 of 26 arms.
Claude Opus 5 + Igala RAG v4.1 93.2 the v4.1 package · 70.5 to 120.4 27 of 43
Raw leak-free chrF 36.8. Rank 5 of 26 arms.
Gemini 3.1 Pro + Igala RAG v2 84.6 the v2 package · 60.1 to 112.2 27 of 43
Raw leak-free chrF 33.4. Rank 6 of 26 arms.
Claude Opus 5 + Igala RAG 83.3 the v1 package · 64.0 to 106.1 27 of 43
Raw leak-free chrF 32.9. Rank 7 of 26 arms.
Gemini 3.1 Pro + Igala RAG 81.6 the v1 package · 65.4 to 100.8 27 of 43
Raw leak-free chrF 32.2. Rank 8 of 26 arms.
GPT-4.1 mini SFT (Igala cold-gold) 40.1 fine-tuned on community answers · 28.6 to 52.8 27 of 43
Raw leak-free chrF 15.8. Rank 19 of 26 arms.
Claude Opus 5 27.2 nothing added · 19.0 to 37.7 27 of 43
Raw leak-free chrF 10.8. Rank 23 of 26 arms.
GPT-4.1 23.4 nothing added · 17.3 to 31.6 27 of 43
Raw leak-free chrF 9.2. Rank 24 of 26 arms.
Llama 3.3 70B Instruct 18.5 nothing added · 11.2 to 29.8 27 of 43
Raw leak-free chrF 7.3. Rank 26 of 26 arms.
Scored on the 27 leak-free frozen questions; the interval says where the score lands 95 times out of 100 if the exam questions were drawn again. Every interval at the top of the board overlaps every other. Only v4.1 over v3 survives a paired test, the audit of September 1, 2026 found. In an arm's name, "Igala RAG" is the packed bundle from one-answer.txt, and the version number is the version from six-versions/.
On the live board, rows marked "tone-stripped" are controls, not systems anyone proposes. They take a finished answer and delete its tone marks before scoring. The September 1, 2026 audit asked for them because the community rarely writes tone marks, so a version that writes fewer of them scores better without speaking any better. A control that sits above the real arms is the audit's point, not a result.
The snapshot holds twelve arms. The live board, when the feed answers, holds every arm that has taken the exam.
The dates are fixed history: what each day added and what it corrected. The entries are copied word for word from the project's own record, which the public method page and the annotation platform share.
Aug 9, 2026 The automatic eval harness, the honest human ceiling, and the leak guard.
The automatic eval harness, the honest human ceiling, and the leak guard. The audit that day found the benchmark had served 15+ of its 43 frozen questions their own community answers - those scores measured copying, so every number since is reported on the leak-free subset.
Aug 12, 2026 corrected sep 1 The Bible parallel corpus, plus a 2,104-entry lexicon, powering retrieval v2 and THE METHOD.
The Bible parallel corpus - 30,907 Igala-English sentence pairs from the Bible Society of Nigeria's Igala Bible - plus a 2,104-entry lexicon, powering retrieval v2 and THE METHOD. Corrected Sep 1: this entry said the pairs were ingested under BSN permission. Our records hold two written requests to the Society and no reply, so no permission is on file. What that means for the corpus is an open item recorded in the Sep 1 entry.
Aug 13, 2026 The frontier arms joined the board.
The frontier arms joined the board. Gemini 3.1 Pro topped it untouched; Claude Opus 5 gained +22 from community retrieval - a clean read on knowledge versus skill. This page was made public and the cost ledger rebuilt.
Aug 14, 2026 A working grammar deduced from all the evidence, and METHOD v3.
A working grammar deduced from all the evidence (tasks/igala-grammar-deduced.md) and METHOD v3, which enshrines only its A- and B-grade rules in the system prompt.
Aug 17, 2026 The benchmark visual and the Community Agreement Score.
The benchmark visual and the Community Agreement Score: leak-free stripped chrF rescaled so the deduplicated native-speaker ceiling reads 100, drawn LLM-benchmark style with confidence whiskers. The raw chrF table moved under the chart; nothing was removed and no score is capped.
Aug 29, 2026 Global Recordings Network signed a copyright agreement (Aug 27).
Global Recordings Network signed a copyright agreement (Aug 27) covering their “Words of Life” Igala recording, and the audio (45:38, the only usable Igala speech asset) was acquired, along with six Bible-for-Children booklets as raw assets; the booklets' fonts silently strip the ẹ/ọ subdots on extraction, so nothing from them may enter the corpus until that is solved. Outreach to other rights holders (the JWAL papers, Egbunu's proverbs study, PanLex) is in progress, with a call with the JWAL author scheduled; none of their text enters the corpus before written permission is on file, so the corpus counters above are unchanged.
Aug 31, 2026 METHOD v4 and v4.1.
METHOD v4 and v4.1. v4 rewrote the instructions around one rule from the community's review: translate the meaning of the whole sentence, never word by word. v4.1 added eight grammar rules mined from 132 judged failures, a step that tells the model to perform a greeting rather than describe one, a rule against inventing dialect facts, and a repair round: when an answer uses letters Igala does not have or is saturated with tone marks, the model is asked once to rewrite it. On the frozen exam Gemini v4.1 scored 120 and v4 102; Claude v4.1 scored 93 against 55 for Claude v3, so the rules that had hurt Claude at v3 no longer do. Nine grammar entries were added to the knowledge store, but the v4 retrieval path does not read them, so they contribute nothing to these scores.
Sep 1, 2026 the audit An adversarial audit of every public number, run against the live database.
An adversarial audit of every public number, run against the live database. What it corrected. A bar past the 100 line is mostly built in: a model is scored against every community answer for a question while a speaker is scored against the other speakers only. Scored like-for-like the best system sits at about 103, level with the speakers, and the score is being replaced by one that cannot pass 100 by construction. The v4 to v4.1 gain is mostly fewer tone marks, which the community rarely writes; with tone marks ignored the two versions are level. The sentence "the grammar lifts Gemini measurably" was not supported and has been removed. The blind preference belongs to one pairing only, Gemini with the v3 package against the same Gemini with nothing added: 54 to 14, with ties and double rejections counted separately, or 25 to 7 when each question is counted once. No speaker has yet judged v4 or v4.1. The fall in "both answers inadequate" from 99% to about half came with a change of models, not from the method. 131 of the 238 exam answers were written after a speaker saw and rejected a model's attempt, so "written before seeing any model" was wrong for more than half of them. And the Bible corpus had been described as used under a BSN permission that our records do not contain. What held: speakers prefer the v3 package to nothing, every one of the six annotators; v4.1 undid the regression v3 caused for Claude; and v4.1 beats v3 on the exam in a paired test, the only step between Gemini versions that does.