Solutions with hashtags have incorrect snippets

In the community solutions tab on the Python track, solution snippets that contain a string with a hashtag (#) in it are messed up by the auto-removal of comments.

For example, this Spiral Matrix solution has an incorrect snippet.

Screenshots of the example


the snippet extractor syntax is very very basic.

I’ve worked with the snippet extractor a couple of times, but I don’t see how we currently match every # except those inside Python strings.

code#comment
foo = '#'

In both snippets, the hashes immediately touch non-whitespace. In the first case, the hash starts a comment and can be discarded. In the second, it occurs inside a string and should be maintained. #\p means match # without requiring whitespace on either side. That gives us code and foo = ' because everything following the matched # gets removed.

This isn’t really a Python-specific issue so the next bit might be the start of a separate thread.


We could add a new modifier that preserves the area between the opening and closing patterns. That’s the opposite of \j which drops the region. Then you could have a line where the snippet extractor sees a starting quote and preserves everything up to and including the closing quote. foo = '#' gets preserved, and that won’t fire on code#comment because there aren’t quotes to match.

2 Likes

Looking through the docs, it does seem like it is an edge case that can’t be covered by the current syntax.

Since such a change would be cross-track (and it would likely be helpful for other tracks too), it makes sense to change this thread to a cross-track one.

Re-categorized this as a general programming bug involving community solutions and the snippet-extractor.

One other place to check # refs gone haywire is the C++ track. IIRC, refs to header files start with a #. But those don’t have a space around the #, which probably makes all the difference.

Here is a copy of a solution in the editor.

Here is one from community solutions. Note that # wasn’t stripped:

One thought is that # followed by letters/numbers and a \n is likely a comment (at least in python), whereas # in almost any other context is not. But that’s probably hard in practice to create/model a rule for.

1 Like

Something like #include would match the include in that example and remove it plus the rest of that line. If whitespace is optional in an include statement, you would need #include\p to remove the whitespace boundary requirement to remove #include<string>. At that point you have a situation fairly similar to the Python issue because "#include" would become ".

With my idea, let’s say the new marker is \k for keep for the lack of originality. "\pk-->>"\pk matches and keeps text between (and including) opening and closing double quotes.
The first \p removes the need for whitespace before the opening double quote.
The first \k starts a “keep” mode where we don’t apply any nested rules until we exit that mode with another \k.
-->> matches everything from the opening double-quote to the closing double-quote. It’s allowed because AFAIK you can’t nest it.
The second \p removes the need for whitespace after the closing double-quote.
The second \k exits our “keep” mode when we hit that second double-quote.

That lets you match and keep "#include" but strip #include ... or #include... if you add #include\p below this rule.

Can anyone tell me, why an Exercism DSL was invented in the first place to solve such a solved problem?

I think awk is an omnipresent tool, that does any kind of text manipulation, is well understood and tested, known to any AI we might ask, fast, and - as GoAWK demonstrates - custom implementations are (AI-)doable if one does not want subprocess spawning?