Rewrite instructions for OCR Numbers

The current instructions for the exercise instructs students to carry out four steps:

  1. Convert “0” or “1” to 0 or 1 (with error handling and “?” for unknown).
  2. Support multiple input digits (with “?” support for a second time).
  3. Extend OCR from “0” and “1” to all digits.
  4. Support multi-line inputs.

However, practice problems aren’t set up to run “steps” and all the test data tests the same “property”/function. As such, I propose rewriting the description to be more direct and not be broken down into steps. (I’ll leave story writing to another discussion/PR/person.)

Proposed description


Optical Character Recognition or OCR is software that converts images of text into machine-readable text. Given a grid of characters representing some digits, convert the grid to a string of digits. If the grid has multiple rows of cells, the rows should be separated in the output with a ",".

Each line of the grid is made of a series of cells, each forming a digit three columns wide and four rows high.” It’d be helpful to indicate there’s no spaces between the digits as well.

  • The grid is made of one of more lines of cells.
  • Each line of the grid is made of one or more cells.
  • A cell is three columns wide and four rows high and represents one digit.
  • Digits are drawn on the grid using pipes ("|"), underscores ("_"), and spaces (" ").

  • The grid is made of 3x4 cells.

Edge cases

  • If the input is not a valid size, your program should indicate there is an error.
  • If the input is the correct size, but a cell is not recognizable, your program should output a "?" for that character.

Examples

The following input is converted to "1234567890".

      _  _     _  _  _  _  _  _  #
    | _| _||_||_ |_   ||_||_|| | # Decimal numbers.
    ||_  _|  | _||_|  ||_| _||_| #
                                 # The fourth line is always blank,

The following input is converted to "123,456,789".

    _  _
  | _| _|
  ||_  _|

    _  _
|_||_ |_
  | _||_|

 _  _  _
  ||_||_|
  ||_| _|

2 Likes

Introduction for `ocr-numbers`? is a post of mine from back in September if anybody want to workshop a story there. It doesn’t have to be my initial idea.

1 Like

It might be best to keep the changes independent but I’m all for that story.

Yup, that was my thinking too. I already happened to have a thread we can direct folks to.

I had similar concerns about Say’s step-based instructions so I’m in favor of updating the instructions.

I think it’s worth specifying the grid cells have three columns and four rows because my brain went three rows and then four columns. “Each line of the grid is made of a series of cells, each forming a digit three columns wide and four rows high.” It’d be helpful to indicate there’s no spaces between the digits as well.

I’d also specify what characters we mean with pipes and underscores. Spaces are pretty self-explanatory and might look weird in the instructions. I was thinking something like “The grid contains pipes ( | ), underscores ( _ ), and spaces.”

Edit - “Each digit in a row is drawn side by side on the grid three characters wide and four lines tall, the bottom line always blank.” avoids discussing cells so that’s one less thing we’re introducing in the instructions.

Thanks for the feedback! I incorporated it into the grid details in my original post.

I find that sentence hard to parse. I think introducing a term is fine so long as it is clearly defined.

1 Like

I like this introduction. My suggestion would be to change:

to

“The grid may contain pipes("|"), underscores("_"), and spaces (” “), but nothing else.”

I would tend to leave the simplest thing here. I would state, if we want to limit:

“The grid may only consist of pipes ("|" ), underscores ("_" ), and spaces (” “).”

Does this suggest the “empty line” since it may also contain an empty string?

Also, it ignores the potential end of line character, if we believe that the examples do not bring out the “nothing else” or other small details that the tests will enforce.

How about, The grid is made up of pipes (“|”), underscores (“_”), and spaces (" ").

This leaves details about newlines, blank lines, etc unspecified. The exact specs and requirements are defined by the tests (and how tracks implement those). The instructions aren’t meant to be comprehensive; they’re meant to explain the overall details.

The grid may contain other characters in the future or on specific tracks; this may have a test for a specific error. The instruction shouldn’t rule that out. The grid may be an empty string. Tests may expect an empty output or an error in that case. The instructions aren’t supposed to dictate those cases; the instructions are merely supposed to explain the shape of the problem.

1 Like

My concern is mainly that a student might interpret the grid contains pipes, underscores and spaces as all valid grids contain at least one pipe, one underscore and one space instead of all valid grids may contain any number of those, possibly none. The and here may be interpreted as AND instead of OR.

I think we shouldn’t focus too much on invalid grids, because they may have any character or form.

Perhaps The grid is made up of pipes (“|”), underscores (“_”), and/or spaces (" ")?

Would it help to focus on the numbers vs the grid?

Numbers are drawn on the grid using pipes ("|"), underscores ("_"), and spaces (" ").

2 Likes

I like this!

With this, and the other constraints, I don’t think there is much ambiguity in this and. We also preserve the possibility of empty and invalid grids.

I don’t have any other suggestions and this seems like a good change, so it’s a +1 from me.

1 Like

+1 from me, too.

1 Like
1 Like

@homersimpsons came up with this comment in our track when I sync’ed the docs to our exercise:

Do we really want to remove trailing blank spaces for input samples in this file?

In the PHP track we had “valid characters” in the code blocks. Maybe somewhere along the journey they got lost in the problem specs? You need to select the code to see the difference in the blocks below.

Our version was this:

 _ 
 _|
|_ 
   

And problem spec does it like this:

 _
 _|
|_

Should the spaces be added (again)? They are present in the canonical test data…

They should be there, they should not be removed.

The whitespaces are semantic, part of the expected structure, if I remember correctly.

Too bad the spaces are not monospaced in the “select to see”

One might even use a ASCII block character to represent the spaces, so that it is more apparent, while having a sample that is actual.

Yeah. The spaces ought to be there. The PR got merged before I could update it. PR to fix the spacing:

prettier complains about the trailing spaces ;) It needs a prettier-ignore pragma :slight_smile:

2 Likes