We're going to translate Exercism to other languages

You already know my overal stance on this which is in agreement with the state of the world and thus with this approach.

I know you dont want deep discussions about implementation do not discuss the technical points I make below :grimacing:

My five cents:

  • strictly separate translation of and contribution to the website ui (yml files) and the course content.
  • please add lang="language code" to the html tag if the website ui has been translated. Keep it on English if you have not.
  • please add lang="language code" to the top-most elements of translated content if different from html tag. This will make screen readers context switch only on that content
  • set Content-Language and read Accept-Language for defaults. Do not forget to add Accept-Language to Vary. Warning: Content-Language: locale does NOT mean the content is in locale. It means the content is meant for speakers of locale. Slight difference but works for us. You can do this the moment you have minimal support for a locale.
  • you should (must) support full locales and not just languages. French french differs enough from Belgian French and Spain Spanish from Mexican Spanish that you want both. Do not underestimate this.

I think using GitHub for translation is fine but you need something visual to show what’s missing. It will be too hard to do if we don’t have that.

4 Likes

A post was split to a new topic: Should we have per-language-tests?

A post was split to a new topic: Dealing with rtl languages

My first feelings are skepticism and fear, as you could expect from me whenever change occurs :slightly_smiling_face:

I don’t think it makes much sense to learn Elixir if you can’t read English well because you won’t be able to read the official documentation or the documentation of any of the popular libraries. If you can’t use those, you can’t build much. If you struggle with reading in English, you should be practicing that as much as possible, not avoiding it.

But maybe that’s a naive approach of an old grumpy spoiled lady. I learned English before I learned programming and had access to good education in both. Maybe nowadays nobody reads the documentation and instead vibe-codes in their native language. No idea.

How many locales are we talking about? If it’s around 50 or more, that would mean I should expect 50 copies of each .md file from my track repo. I am wondering whether that drastic increase in the repo size will have noticeable performance impact on git commands or my IDE’s indexing of the repo.

I agree that LLMs are “good enough” at translating, at least based on my own experiences with a very popular translation direction of EN ↔ DE. I also agree that having everything in Github would be most convenient for me as a maintainer, if it doesn’t cause significant performance issues. It will give us a lot of flexibility to develop automations in forms of scripts and Github actions, and I’ll have a lot of visibility into how those automations work.

7 Likes

Everything Angelika said :slightly_smiling_face: (she is much better than I at articulating these things!)

And in CPython’s case, and extra “FU”, because the language itself does not have a mechanism for translating any stack traces or error messages. You have to build your own distro for that.


…and 50+ copies of any instruction appends and approaches docs sets…


My first thought here is — what happens when an instruction append file is added where there wasn’t one before? How will we notify/track incoming translations? And what about approaches articles? There is probably more I am forgetting…

But the hope would be that we would have some sort of configlet command to surface what’s needed/missing … or something…

Edited to add : Maybe also trying to think through what the “pipeline” / timing looks like for adding new practice and concept exercises. I am assuming the “base” add looks the same as it does now … but then we have the tracking of all the translations.

2 Likes

I miss our community calls :blue_heart:


So I appreciate you and @Meatball speaking up. This is my rationale…

Around 20-25% of people can read english to a “useful” extent (according to random internet sources). So we’re left with 75-80% of people who can’t. I’ve been in Japan for the three month and used Google Translate on pages daily, and it’s sort of ok, but often confusing, and very clunky to use. Finding well translated pages is a real joy. I imagine the experience is the same in reverse for Japanese people, etc.

While this is true, I imagine a lot of people are “coping” by browser-translating these, which is maybe of for docs, but for education feels like a huge extra barrier to overcome.

As those of you who have been around for a while know, I’ve generally been skeptical of this whole idea, but my skeptism has come from the technical barriers rather than the “is this the right thing to do” perspective. I think with LLMs to help, this feels like the right time, and if a primary goal of our organisation is helping those who don’t have the resources to learn elsewhere, I feel this offers an opportunity to help a vast amount of people.

This has also been by far the most requested feature we’ve ever had, which I don’t feel would be true if it was a “solved” thing with browser-translation.

Also, doing the Bootcamp, we had very few people who didn’t speak English as a first language, and no-one (I know of) who wasn’t fluent. I think that’s indicative of the need for resources in other languages. (or they’re all covered already - but that’s not my understanding)

I feel like the work for maintainers is low. Iit’s lots for me but the tradeoff is worthwhile in my “the point of my work is to help people” mindset. And for the people who it will cause work for (those who want to translate), it’s work they want to do. So I feel the general burden is ok.

So that’s why I came to the conclusion this is the right time to do that. But I’m always appreciative of counter-arguments, especially before I sink a ton of time into it!

1 Like

My thinking is it gets immediately auto-translated by LLMs then can be “touched up” manually in a subsequent “Update this for languages” PR, which stays open for ~x days before merging (to avoid a million little PRs for you). And I’m also imagining we have a “sanity-check” GHA that checks the translations are correct and someone hasn’t inserted something that we can’t review easily.

But we definitely need subsequent discussions about this sort of thing.


In terms of repo size, my instinct is that even 50xing the size of a repo still won’t be that big compared to the codebases we tend to work with day-to-day. But we should do maths to check that!

I raise you several decades on being old and grumpy! The manual, in a well-thumbed hardback copy:

(But with nothing useful to add, I’ll shut up now)

Another thought here about marking/notification. What happens if we never get a human translator double-check? I think maybe we note where something was machine translated and then not reviewed vs translated and reviewed.

Two more thoughts on process:

  1. The (way not fun) process of finding/reporting typos once something is published. Where do they go, and how do we deal?

  2. What happens when translators disagree on the type/specificity of the instructions, and add in more specification / different instructions? How do we check “good deviation” from “bad deviation”? Or do we not worry?

2 Likes

I see a lot of valid points and concerns. Here’s my 2 cents.

  1. I think it is very hard for a lot of us to appreciate how much of a barrier the English language can be, seeing as we’re all discussing this (semi-)fluently in English here :slight_smile: If I try to put myself in the shoes of someone that struggles with English, I can imagine having native translations being incredibly useful. The majority of books I was given to teach me how to programming were in Dutch, not in English.

  2. One problem of the built-in browser translation approach is that it doesn’t really work for CLI-based users. Yes, they could have the browser open to read the instructions, but that is less ergonomic.

  3. Translations do not have to be perfect to be a (huge) improvement over the existing situation.

  4. Having the translations in their own sub-directory and not directly next to the existing docs seems like a good step to somewhat reduce the “noise” in the repo.

  5. I think we should start out small with just one or two languages and a couple of exercises, just to get a feel for the issues and workflow. Personally, I find it hard to estimate how much I will like/dislike the maintenance part of it without actually seeing how it works (I’m not very imaginative…)

5 Likes

Translations do not have to be perfect to be a (huge) improvement over the existing situation.

My concern here is how much better over browsers translation will this project result in. I find google builtin translator to be pretty good, there are certain things it does translate incorrectly, like the language names and codeblocks isnt translated. But codeblocks is anyways quite tricky to translate since if you start applying it to a llm that can quite quickly start to translate actual code. Furthermore say the maintainer or maintainers of a language just disappears (or don’t have time to contribute anymore), that language will just be auto translated. Will that llm be better/on par with google translate or will it be worst? And I would just guess but there is likely thousand+ documents. Translating that to 50 languages, that is 50k documents, just to me feels a bit unfeasible. Even if it is just 10 languages, 10k pages is a ton.

One problem of the built-in browser translation approach is that it doesn’t really work for CLI-based users. Yes, they could have the browser open to read the instructions, but that is less ergonomic.

I agree, here is somewhere that translation could make sense. But how many users actually use the cli? If there isnt a very large user group, perhaps a 100% automatic system with using llm and having some disclaimer in the beginning that this is machine translated and might have errors.

I might add that I am on Angelika side that I think English is a fundamental skill to have. My culture is quite English centered, they start teaching English at an age of 7-8, and the majority of entertainment that is not kids focused is in English. You can only find kids movies dubbed, there is very few video games available in my native language. This lead me to be able to speak quite fluent English (according to myself) at an age of 13-14. But I also know that if you don’t get exposed to a language you wont learn it, I studied French for 4 years, now 4 years later. I know how to say what my name is, how old I am and hi. I know some more basic expression, but my knowledge is more or less non existent these days and I were never really good at it.

I have no good experience with auto-translating technical stuff from English to my native language. The main problem is, that LLMs tend to translate much too much. This always makes me switch back to English whenever there is an automated translation presented to me. It’s definitely not “good enough”.

Also there is a fine line between helping to learn programming and preventing to learn the right words and phrases to go out and find things on the internet. For most of the more complex topics, knowing the correct english word to find answers to questions is key.

So, while I do support helping people in their native language to learn programming, this should not take away all the required learning of English and tech vocabulary.

We have the exercism-cli, it can connect to an api for translation, and there are terminal based translation tool. Even from text to speech.

So I would say that for the terminal users, they download the exercise, they have the tool available, we can have the exercism-cli tool as the entry point to the wanted translations.

I have used trans, for example.

When new students fail with “Hello, World!” we tell them that Exercism is not meant for people without prior programming knowledge. The BootCamp changes that aspect of exercism.

I have the feeling that the biggest bunch of people, who are not fluent enough in English to use the website, are beginners and would most benefit from a translation of the BootCamp.

I have no data at all to support my claims, but I cannot imagine anyone without decent knowledge of the English language to navigate a language like C++ or UIUA to a point, where they can solve harder exercises.

So I would vote for two things:

  • make the BootCamp accessible first
  • keep the i18n files in a separate directory, so it is not cluttering up the repository. Expecially if it is the same exercise text for every track. (Many of these are verbatim copies across tracks anyhow, right?)
2 Likes

I appreciate everyone’s thoughts.

I posted on LinkedIn to see if we can get some wider opinions. Any reposts for you networks would be appreciated: Should we translate Exercism into your language? | Jeremy Walker?

I’m acutely aware that our maintainers are some of the most capable and smart people I’ve worked with, and also that most learnt some basic English quite early in life. Its hard to know if your experiences would match someone with a lower efficacy or who are starting from zero with their English later in life (both of which are most people in the world).

1 Like

Disclaimer: The following is my initial opinion on the matter, not having considered many of the more technical aspects.

I think I have some mixed feelings about this, but most of them will probably be changed as this project gets rolling.


The majority of the people I know personally are multilingual and quite fluent in English, so the language barrier doesn’t seem that big of a gap to me personally, but It definitely exists. It’s also somewhat sad that good learning resources (for programming) in my native language don’t exist because everyone seems to assume that English is ubiquitous.

In support of translations, I can provide a personal anecdote about my dad. He’s oldschool, he has studied English at a later point of his life and isn’t quite fluent but has good enough understanding to deal with most day-to-day things. Recently he has expressed an interest in learning some higher level programming, but the majority of the resources are in English and it’s been enough of a hurdle to discourage him, so I can see how having at least some translation would help, regardless if it’s machine or man made.

That being said, I support the point that was made that English is somewhat necessary for learning programming and specifically technical English terms like currying or short-circuiting or hoisting etc. Translating things like this wouldn’t necessarily add much value, because as @mk-mxp said:


I think the translation problem/project can be separated into two topics:

  1. Translating the site, the docs and the learning resources
  2. Translations during mentoring

About the first point, I’m apprehensive because having 50+ locales and 50+ versions of every .md file in the repo of a given track would add a lot of maintenance overhead, or at the very least clutter things up a lot. It would also make it harder to maintain translations for one language/locale, because you’d have to work across all track repos. So from that point of view, I think a central separate repo for each locale that contains all the translations for that locale/language would be easier to maintain.

Also, what happens if a certain language/locale doesn’t have maintainers? Do we just go with the “good enough” machine translations and add a disclaimer? Do we deprecate the locale when it is unmaintained for X amount of time?

About point two, I often see mentoring requests from people in their native languages and those requests typically stay for a bit longer in the queue, so i think that having translations there would be most beneficial. It also makes the signature unique feature of Exercism more accessible. Even if it’s some sort of ad-hoc machine translation that isn’t stored anywhere, it would be easier to communicate things like this, so I’m all for this point.


Lastly, for now, I know you said you don’t want to discuss the technical aspect of things but I can’t help but ask, why not consider some ready made localization tool to help, instead of building our own or relying entirely on GitHub. Maybe something like Mozilla’s pontoon or similar?

I also think it would be too hard to get this project going if we don’t have something to help keep track of translations and visualise things.

2 Likes

Hi there,

I was reading about translating Exercism and I think I don’t have much to add concerning the technical part of the task (huge task). About the usefulness of the content in other languages though, I can say that in places like Brazil (where I live) this would also be an action to help the less fortunate since the majority of people who speak English come from an advantageous background compared to those who did not have the opportunity to learn another language (I myself only learned English well after my 30s). So, those people are left off from having access to amazing content when they try to learn how to code. BUT, I think English is a must to have for anyone who wants to work in tech. So, maybe having the content translated would be like a scaffold only until the person gets better in English and in coding. I know the freecodecamp did this a while ago and now they have even an English course for programmers among the programming content.

Anyway, if you decide to go through this endeavor I can help to revise Portuguese. Just let me know.

Great platform by the way.
Thank you!

6 Likes

I’ve posted on discord about this, but as a fellow brazilian I have the same opinion as Sabrina. And I can also help with translating to (brazilian) portuguese.

2 Likes

A broader approach in raku is an ongoing effort at localization to translate code so that you can code in your native language and be supported by someone whom cannot speak your language by having keywords automatically translated into a language they can understand.

For example the hello-world exercise could be written by a German learner as:

sag 'Hallo Welt'

and automatically shown to an English speaking mentor as:

say 'Hallo Welt'

using the lookups in DE.l10n.

Human language translation is a worthwhile and problematic effort. Computer language localization could be another aspect to consider for languages which support it as it appears a less complex undertaking.

1 Like

I’ve shared my thoughts in discord and you’ve already raised a lot of good points here, so I’ll keep it short:

  • Experienced engineers or eloquent English speakers should be able to handle non-translated resources well already, so translations would be most beneficial to 1) young people, 2) complete beginners (who may kindly be redirected towards the Bootcamp, if possible), 3) people overwhelmed with algorithmic thinking who would rather focus on that part.
  • Thus, before focusing on the exercises, it may be better to focus on the website and bootcamp.
  • Locales are important, but reaching out to major language groups based on popularity may be more beneficial (probably requires a bit of factual research).
  • I still have nightmares from old courses using #typedef int akeraios (ακέραιος being integer), I consider it a mental handicap actively doing harm to the learning process.
5 Likes