Hi everyone ![]()
Over the second half of 2025, we’re going to add internationalization to Exercism with the aim of supporting as many other languages as possible ![]()
This is a preliminary post to help me think through the project before we get started. It’s really aimed at maintainers or contributors who understand the Exercism ecosystem well. My hope is that you’ll help me understand all the things I’m missing in my thinking! ![]()
Philosophy
We are translating the site because we want to help people who use a natural language other than English. We are striving for a baseline of “good enough”, and a peak of “native-level”. LLMs are now at the stage that they give “good enough”. The translated language might be chunky or messy, but from what I understand, it is understandable.
I do not, therefore, want us to concern ourselves overly with enforcing rules that ensure perfect translation, but accept that just the act of doing this will hugely benefit people, and “good enough” is a good starting point. However, we do have a huge community, many of whom are passionate about good translation, and if those people want to help by investing their time in making more natural translations, we should make the most of that.
I am also aiming to minimise any increased burden on maintainers (and myself) to achieve this.
Overarching Structural Decisions
There are three overarching decisions that fundamentally change how this project works. I’m pretty set on these, but if there’s some consensus that this is terrible, I’m open to hear it.
- The most fundamental decision is that I think we should store translated text in GitHub, in the same repo as the content. For example, website translations will live in the website. Exercise translations will live in track repos, and in prob specs etc.
- We will start by translating everything using LLMs. People can then fix the translations for their language. (We’ll probably have a 2-3 month project where we get good-first-versions of everything, then merge those PRs in bulk - rather than having thousands of small PRs)
- In the future, when there is a change to an English markdown file, we will use an LLM to automatically update all the translations, only touching the bit that’s changed in the English version. This means our quality of translation should only degrade in small chunks, and can be checked/patched by a language team, without having to re-review the whole file.
I have considered doing this outside of GitHub via a custom UI in the website, but I think it’ll just make the whole thing lots harder to manage and give much less visibility - and also take a lot longer to make.
The Technical Mechanisms
Naturally as developers, we tend to nerd-snipe ourselves into thinking of the low-level (interesting) bits of this. I’d like to not get too caught up in that in this discussion, as I’ll put together dedicated discussions for this later.
In terms of the repo structure, I’ve not really thought about it yet. We’ll need to balance how things like the CLI works with the nice structure of the exercises. I suspect having an i18n folder in each exercise in a track repo, with introduction.nl.md etc in would work well. And something similar in problem specifications. The website bit is pretty complex, but you all don’t really need to care about that.
In terms of keeping things updated, my general thinking is to use GitHub Actions, which can be triggered by maintainers, to schedule an auto-translation and add a commit to any PR with that translation in. I think we’ll then have another GHA that checks that any time an english file changes and ensures all other files have been updated before allowing a merge.
We’ll have space to have discussions on these points in the future, but I just wanted to seed ideas for you in this post.
People
I’m imagining we will want 1+ people for each natural language to be responsible for monitoring and fixing translations. I don’t know how to manage this side at all tbh, but we can probably create a “exercism/languages” team that has readonly access to all repos, can get pinged when things changed (via the GHA bot) and then can sign things off in comments or checkboxes or something on the PR. I’ll think about this when I’m a little less burnt out!
Mentoring
I intend for people to be allowed to specify their preferred language during mentoring, but mentoring to occur in English still, with translations shown. Again, pragmatically, I think this will help lots of people who currently wouldn’t use our most beneficial feature. But also I see this as quite experimental, so intend to mark it as such.
I’d appreciate any comments on:
- Emotions: Joy/encouragement/terror/horror/etc.
- High-level problems people foresee.
- Systematic/structural thoughts.
I’d rather we avoid the lower level implementation discussions for now, else this thread will become unwieldily quickly. If people start getting deep on tangents, I’ll reserve the right to just delete those posts from here, so please respect that! ![]()
I’ll be posting a full announcement video and blog post sometime in the next few weeks about this, once maintainers have had a chance to comment and I’ve had a little post-bootcamp/bots rest!
Thanks!
