Erlang numbers their function names so there’s no duplication. Your example test becomes 7_does_not_detect_non_anagrams_with_identical_checksum_test_ but is reported by the runner with a test name of “does not detect non-anagrams with identical checksum”.
The UUID thing is a bit problematic IMHO. These files also get downloaded by students using the CLI, and it feels pretty confusing/noisy to me to have to look for what the test function is testing when scanning the file (especially if the files are inconsistent in their formatting). But I am also brainwashed by Python’s “readability” mandates.
Leaving the naming aside, I’m chiming in for @BNAndras method/suggestion.
Python has a test generator, and uses a JinJa template per practice exercise for all but a very few exercises (here is an example of the JinJa template for Anagram).
This avoids the duplication and errors from manually updating test cases for the most part, since generation is based on a tests.toml file that lists test cases by UUID. The only (potential) hiccup is when a bunch of test cases get reimplemented, and we are sloppy in updating the tests.toml file.
We pull the canonical data description as the test function name, and because we’re using the syntax from Python’s built-in unittest module, we have both a class name and test functions that all start with test_ (here’s Anagrams test file as an example).
This led to extremely long test names and weird wrapping when we first tested it in the V3 UI, so we had the test runner trim and reformat it for the test runner JSON. Here’s the code and here is what that looks like on the site:
What the test function is testing is indicated by the (testing "Does not detect non-anagrams with identical checksum") part. That’s exactly why it exists. Function names are not supposed to used as a substitute for test descriptions.
Compare the two versions:
(deftest does-not-detect-non-anagrams-with-identical-checksum
(testing "Does not detect non-anagrams with identical checksum"
(is (= [] (anagram/anagrams-for "mass" ["last"])))))
(deftest test-1d0ab8aa-362f-49b7-9902-3d0c668d557b
(testing "Does not detect non-anagrams with identical checksum"
(is (= [] (anagram/anagrams-for "mass" ["last"])))))
In the first version you first parse the function name. Then you read the actual test description. You’ve just wasted time parsing two identical things formatted differently. In the second version you still read the first part of the function name, but you make no effort to parse it. You already know that the test description is on the second line.
I was of the same opinion, and I definitely haven’t been brainwashed by Python’s reqadability mandates
I honestly don’t see the problem here. Once I’ve seen the first version (with the longer name), any subsequent reads will just make me skip that bit (same as the guid bit).
To me, it looks like we’re optimizing things for the maintainer, not the student, whereas it should really be the other way around.
It’s a combination of both. Including the description in both places isn’t common practice, and I prefer not to imply that this is how tests are typically written. The (testing ...) part should describe the test, while the function name can be anything. Ideally, it would be a concise description, but that often leads to duplicate names. Moreover, having the description in both places isn’t ideal for user experience, even if someone trains themselves to ignore the function name.
That said, I now have a clearer plan for moving forward. I’ll share my decision here later, but first, I’d like to gather a bit more feedback if possible.
Reading some example code, deftest is often defined once per tested function (e.g. deftest isogram?. IIRC the downside of that was that it’s inner tests stop executing when the first error occurs?
This exercise has all tests in a single deftest. I edited the circled code so that more than one test would fail. (I added a zero to the end of each number 1->10, 2->20. and so on)
Yet, only the first failing test case is shown. The “ones”.
Edit: This appears to be a test runner issue. Can’t replicate locally.
But, if only the first failing test is shown, it might also mean that the inner tests never execute. So i guess the correct reply would be "No, they should execute properly. "
Yeah, property identifies what’s being tested so it’s more or less synonymous with the function name in practice. You’re not held to using that name though as represented especially if it’s not idiomatic to your language.
Alright, I’ll wrap things up. Thanks to everyone who viewed and shared their thoughts. Here’s the plan moving forward. Everything below this line is Clojure-specific:
If the test runner is fixed to display all failing cases, we can consider:
Test Organization:
One deftest per tested function will include all test cases for that function, along with their descriptions. The name of the function will follow the test-<function-name> pattern, where <function-name> matches the name of the function in the stub. For example: test-anagram?.
UUID:
The UUID of each test case will be included as a comment (e.g., ;; <uuid>) before each (testing ...) form so that we can quickly locate each case in the canonical data.
Downside:
This approach increases code density within a single function. It can become unwieldy, especially when implemented tests span 10+ lines. The inherent nesting of Lisp syntax exacerbates this, making it harder for humans to parse compared to having one test case per deftest.
If the test runner isn’t fixed:
Test Organization:
One deftest per test case. Each test case will include its own description. Each deftest will be named using the corresponding description as it appears in the test file. Given that we’ll end up with the same information in both the description and the function name, the name of the deftest will probably be revised. A possible solution would be the <function>-test-<n> pattern, where <function> matches the function name in the stub, and <n> is a number. For example: anagram?-test-1.
UUID:
The UUID of each test case will be included as a comment (e.g., ;; <uuid>) before each deftest form
Descriptions:
I’d prefer that the .toml file keeps the description verbatim from the canonical data. However, the description in the implementation may be modified. For example, if the description references lists but the implementation uses vectors, the implementation’s description should refer to vectors.
Function Name Shown in the Online Editor
We cannot generate them from the .toml descriptions. I’ve encountered many case where the .toml files were in sync, but the tests have not been implemented.
We cannot generate them from the (testing...) forms because not every implementation has one.
Bummer!
Next Steps:
If the test runner isn’t fixed, I will update all merged and unmerged cases, remove the uuids from the function names, and probably go with the <function>-test-<n> pattern since i don’t see any compelling reason to have the test case description duplicated in the function name.
(deftest largest-series-tests-pass
(testing "can find the largest product of 2 with numbers in order"
(is (= 72 72)))
(testing "can find the largest product of 2"
(is (= 48 48)))
(testing "finds the largest product if span equals length"
(is (= 18 18))))
to this JSON:
{
"name" : "largest-series-tests-pass",
"status" : "pass",
"test_code" : "(testing \"can find the largest product of 2 with numbers in order\" (is (= 72 72)))\n(testing \"can find the largest product of 2\" (is (= 48 48)))\n(testing \"finds the largest product if span equals length\" (is (= 18 18)))"
}
I’ve looked into this and it looks like it might not be possible to fix the test runner, so my preference would be to go with having one deftest per test case and I’m fine with the <function>-test-<n> pattern.