New Inventory Management test(s) didn't run on past solutions?

When looking through the community solutions for Inventory Management, I came across multiple solutions that said “Passed”, but when I tried the code myself, it did not pass all of the tests. Both solutions linked above seem to fail the “decrement items not in inventory” test).

I thought that solutions normally have the new tests run on them when the tests are updated… Is there some reason that didn’t happen on Inventory Management, or is this a bug?

This is controlled on a per PR basis. The maintainer needs to take into account the specifics of the change and the cost of running all the tests. Sometimes tests are rerun, other times they are not. It sounds like this might be a test change that was not accompanied but a rerun.

I see. I was rather confused by this, as I was looking through the community solutions to see if there was more concise solution for that task, but when I thought I found one, it turned out to not take the aforementioned case into account.

I think it would be less confusing if there was some kind of indicator that the version of the tests the solution passed was an old one in that case, but I suppose that’s a different topic.

There’s an indicator on the solution page itself when a student’s solution is pinned to an older set of exercise files. See johnburkert's solution for Jedlik's Toys in C# on Exercism for example.

We have exemplar solutions for each concept exercise on the track and when a new test is added we test that solution to see if it still passes.

If it continues to pass and the exercise has fair amount of completions (this one has about 18,500 right now), we typically don’t “invalidate” it (mark all solutions to be retested). It’s expensive to do so, and likely to not find that many failures. We’d love to always retest everything, but we run on donations and can’t really afford to retest every time on every track.

We do this on a case-by-case basis. In the cases where the exemplar solution doesn’t pass, we know for sure that requirements have changed, and we typically then retest everything. We also retest if we’ve missed a significant case in the original tests, or if we feel that most solutions would fail the tests.

At the time the new test was introduced, we made a call to not retest.

Unfortunately, that means some of the examples you viewed were invalid for the current tests.

There’s an indicator on the solution page itself when a student’s solution is pinned to an older set of exercise files. See johnburkert’s solution for Jedlik’s Toys in C# on Exercism for example.

I think that only comes up when the invalidator is set off, and I think that’s tied to the presence/absence of [no important files changed]. But I could be wrong about that.

I see, it makes sense that the tests weren’t rerun then. It’s odd that the “outdated” indicator doesn’t appear sometimes, though. If it was there, I probably wouldn’t have gotten confused.

I imagine it can’t be that horribly hard to mark students’ solutions as old even without an invalidator present on a technical level. A student’s files are pinned to a particular set of exercise files so the site might be able to compare the pinned files to the current ones. If they don’t match, the student’s files are likely old.

However, that might be confusing in the UI. How do we distinguish old solutions from out/of-date solutions simultaneously? They represent different things so I don’t know what could be done here.

Perhaps there could be an “:warning: Old” indicator with a message that appears on hover that says something like “This solution passed an old version of the tests. It may be outdated now.”

Since the “:warning: Outdated” indicator is more definitive than that, if both would be present there would probably be no need to show the “:warning: Old” indicator.

Many test changes wouldn’t invalidate most solutions, though. Does it make sense to mark all solutions as old when they may very well be perfectly fine?

1 Like

True, but there not being any indicator that the tests were updated since then seems to be a cause for confusion. Maybe it would work if the “old” indicator was less prominent than the “outdated” one? Or perhaps there’s not really a good solution for this…

The problem is that updating the tests can mean a change that has very very little impact to a change that has significant impact. The change can have impact on some solutions and not others. With unlimited resources, we could rerun the tests and determine if a solution is or is not impacted by the change. When rerunning all the tests has a significant cost, it’s hard to determine the impact and any guess made for all solutions will likely be wrong for some.