In the community solutions tab on the Python track, solution snippets that contain a string with a hashtag (#) in it are messed up by the auto-removal of comments.
For example, this Spiral Matrix solution has an incorrect snippet.
In the community solutions tab on the Python track, solution snippets that contain a string with a hashtag (#) in it are messed up by the auto-removal of comments.
For example, this Spiral Matrix solution has an incorrect snippet.
I’ve worked with the snippet extractor a couple of times, but I don’t see how we currently match every # except those inside Python strings.
code#comment
foo = '#'
In both snippets, the hashes immediately touch non-whitespace. In the first case, the hash starts a comment and can be discarded. In the second, it occurs inside a string and should be maintained. #\p means match # without requiring whitespace on either side. That gives us code and foo = ' because everything following the matched # gets removed.
This isn’t really a Python-specific issue so the next bit might be the start of a separate thread.
We could add a new modifier that preserves the area between the opening and closing patterns. That’s the opposite of \j which drops the region. Then you could have a line where the snippet extractor sees a starting quote and preserves everything up to and including the closing quote. foo = '#' gets preserved, and that won’t fire on code#comment because there aren’t quotes to match.
Looking through the docs, it does seem like it is an edge case that can’t be covered by the current syntax.
Since such a change would be cross-track (and it would likely be helpful for other tracks too), it makes sense to change this thread to a cross-track one.
Re-categorized this as a general programming bug involving community solutions and the snippet-extractor.
One other place to check # refs gone haywire is the C++ track. IIRC, refs to header files start with a #. But those don’t have a space around the #, which probably makes all the difference.
Here is a copy of a solution in the editor.
Here is one from community solutions. Note that # wasn’t stripped:
One thought is that # followed by letters/numbers and a \n is likely a comment (at least in python), whereas # in almost any other context is not. But that’s probably hard in practice to create/model a rule for.
Something like #include would match the include in that example and remove it plus the rest of that line. If whitespace is optional in an include statement, you would need #include\p to remove the whitespace boundary requirement to remove #include<string>. At that point you have a situation fairly similar to the Python issue because "#include" would become ".
With my idea, let’s say the new marker is \k for keep for the lack of originality. "\pk-->>"\pk matches and keeps text between (and including) opening and closing double quotes.
The first \p removes the need for whitespace before the opening double quote.
The first \k starts a “keep” mode where we don’t apply any nested rules until we exit that mode with another \k.
-->> matches everything from the opening double-quote to the closing double-quote. It’s allowed because AFAIK you can’t nest it.
The second \p removes the need for whitespace after the closing double-quote.
The second \k exits our “keep” mode when we hit that second double-quote.
That lets you match and keep "#include" but strip #include ... or #include... if you add #include\p below this rule.
Can anyone tell me, why an Exercism DSL was invented in the first place to solve such a solved problem?
I think awk is an omnipresent tool, that does any kind of text manipulation, is well understood and tested, known to any AI we might ask, fast, and - as GoAWK demonstrates - custom implementations are (AI-)doable if one does not want subprocess spawning?