The x86-64 assembly syllabus is live

I meant what keiraville mentioned above.

In general, if you want to clear a register without any further semantics (maybe it is just a counter, for example), then xor is a common choice.

And if the clearing has some larger meaning, such as FALSE or EMPTY or VALID, etc, then it is more natural to have a equ constant and use mov.

As for why xor, it is as keiraville said.

Everything is data, code is just executable data. Every instruction is encoded in a certain number of bytes in section .text. Whenever you use an immediate, it in principle occupies space in section .text. So mov rax, 0 occupies more space than xor rax, rax because all bytes for 0 are actually placed in the code.

But this can be optimized away by the compiler.

And if you are clearing memory, you’d use mov because you can’t do xor [rax], [rax].

For bulk store/move to memory, there are some string instructions that are concise and fast (they are explored in the strings concept). Hint: your intuition about rcx is on point.

1 Like

Strings (Poetry Club Door Policy)

Did you know that I originally wrote this exercise? Story et al! Hopefully this will help me today.

I do my regular stuff with one change: I saw in a community solution you can do multiple global exports in one line:

global front_door_response, front_door_password, back_door_response, back_door_password

I can now run the tests and two pass without any implementation:

I also see Expected 83, Was 0. I think this means that my expectation of characters just being bytes, and likely being ASCII (aka there is no encoding) is correct. I do think it would be helpful to see in the error message both the byte and printable character (if any).

Gotta start reading!

It is possible to declare a string in NASM using the directive db , as usual for declared data

The usage of db reads as if it is special, but I don’t believe it is. I think I already used it before. I think it means “byte” size entries. Perhaps could use a tiny mention that this is just the same as with the arrays we learnt about before. To me, as “usual” doesn’t convey “it’s not special”, especially because of the next part:

Even though this defines an array, it is not necessary to use a comma (, ) to separate values:

Is this also the case for non-string arrays? I don’t know but I continue.

Reading the instructions further, I finally realise that d as in dword probably means double and q as in qword means quadruple. Maybe this was mentioned somewhere, but I do not recall.

As I read the string instructions (lods etc), I do not yet understand how to use them or why to use them, but abstractly I think I need this to copy characters out of the array or copy (parts of) arrays. We’ll see. Oh and yeah, again rcx being special! The note reads weirdly:

Note, however, that rcx is not an operand to those instructions. Its value must be adjusted before using them.

I think better would be something like:

rcx is the register rcx. Just like with loop instructions, you cannot set the value to rcx by an operand to the instructions, but it must be set before the instruction is used.

front_door_response

I need to get the first character in the string. The string is “an array” of bytes, and what’s passed in is the address to the first sentence (array/string). That means that that the first byte at that address is exactly what I need.

Let me first try to mov the first byte. I’d like to try out something like movs, but I find the instructions pretty unclear. The summary says it will copy from rsi to rdi, but the explanation earlier says it will “copy 1 byte, changing rsi and rdi by 1” and even earlier " They usually expect the source to be a memory location in rsi and/or the destination to be a memory location in rdi ."

The summary and first description tell me: rsi and rdi need to be memory locations and then it will copy 1 byte from left to right (or in reverse given the DF). So I don’t understand what this “changing x and y by 1” is. Is it automatically changing the memory addresses by 1 so you can keep copying? Unclear.

There are also no examples. So mov it is!

mov al, [rdi]
ret

So that worked. Next!

front_door_password

It must modify this string in-place, making it correctly capitalized.

Capitalization of ASCII characters is deceptively simple. For example the letter A is represented as 0x41 (or 65) and a is 0x61 (or 97). They are 32 apart. Capitalizing a letter is lowering its byte value by 32 and doesn’t need to happen if the value is already smaller than 97

So if I wanted to titleize a word (not uppercase), it would be:

mov al, [rdi] ; temporarily store a copy of the first character
cmp al, 0x61  ; check if already capitalized
jb .skip      ; skip capitalization if already capitalized
  
sub al, 32    ; capitalize
mov [rdi], al ; write the character back to the memory
  
.skip:
ret

However, I want to explore how I could do this for each character. There is no length given for the string, but the exercise states each string is terminated by \0. That can be used to loop and break. (I’ll probably need this later anyway, so let’s try).

I think I can keep a counter to increase the memory address and break when the \0 is reached. I want to use loopne, but I also don’t want to give an arbitary size. I am wondering if I could set rcx to -1 and have it loop until loopne. Not going to try tho, I don’t need to. jne gives me that behaviour.

.upcase:
    call upcase_one      ; (there is no empty password)
    inc rdi              ; move memory by 1 byte
    cmp byte [rdi], `\0` ; check the next bye
    jne .upcase
    
    ret

upcase_one:
  mov al, [rdi]
  cmp al, 0x61     ; check if already uppercase
  jb .skip         ; skip capitalization if already uppercase

  sub al, 32       ; upcase
  mov [rdi], al    ; write the character back to the memory

  .skip:
  ret

I tried it out and it works!

Expected ‘Summer’ Was ‘SUMMER’. Combined letters are: SUMMER
Expected ‘Sophia’ Was ‘SOPHIA’. Combined letters are: sophia
Expected ‘Code’ Was ‘CODE’. Combined letters are: Code

Ah okay. These cases indicate that I could of course get capital letters at non-zero positions.

I need to:

  • upcase the first letter (uppercase)
  • downcase everything else (lowercase)

That should be easy now:

downcase_one:
  mov al, [rdi]
  cmp al, 0x61     ; check if already lower case
  jge .skip         ; skip capitalization if already lowercase

  add al, 32       ; downcase
  mov [rdi], al    ; write the character back to the memory

  .skip:
  ret

Then combine it:

   call upcase_one      ; (there is no empty password)
   inc rdi              ; move memory by 1 byte
   cmp byte [rdi], `\0` ; check the next bye
   jne .downcase        ; break if last character
   ret  

.downcase:
    call downcase_one      
    inc rdi              ; move memory by 1 byte
    cmp byte [rdi], `\0` ; check the next bye
    jne .downcase
    
    ret

image

Rejoice!

back_door_response

What I can do here is:

  • store the current byte if it is a printable character
  • go to the next byte
  • break if \0
.store:
    cmp byte [rdi], 0x41 ; check if uppercase A
    jb .next
    cmp byte [rdi], 0x7A ; check if lowercase z
    jg .next

    mov al, [rdi]        ; store as return value
.next:
    inc rdi              ; move memory by 1 byte
    cmp byte [rdi], `\0` ; check the next bye
    jne .store
    
    ret

image

I feel very good at my process trying to complete this.

back_door_password

I need to do two things:

  • format the string
  • add ", please."

My process will be:

  1. copy the second argument’s memory to the first location
  2. call front door password to capitalize it
  3. “add” the second string at the end (dunno how yet)

I re-read the extra instructions:

instruction description
lods loads from the memory location in rsi into rax
stos stores rax into the memory location in rdi
  • rsi happens to be the second argument passed
  • rdi happens to be the first argument passed

So if I want to copy the string from rsi into the rdi location, I think I can do it by calling lods folowing by stos. However, I think this is what movs is supposed to be for: “copies between memory locations (rsi to rdi)”.

I’ll just have to deal with that “memory increase thing”.

I don’t think I can use rep here because I do not know the length of the string and I don’t think I should use repe or repne because cmps also moves “the pointers” and idk, feels weird. Let me just try to do it manually:

  lea r8, [rdi]        ; store address of destination so I can "rewind"
.move:
  ; lodsb
  ; stosb
  movsb                ; move one byte from rsi to rdi, inc. both
  cmp byte [rsi], `\0` ; check if next byte rsi is end of string
  jne .move

I will try this out first, because the test output will tell me if this is correct or not.

Expected ‘Work, please.’ Was ‘work\xFC\x7F’. Combined letters are: work
Expected ‘Horse, please.’ Was ‘horse\x7F’. Combined letters are: horse

\x7C is the delete (special) character, and \xFC I think is in the extended set? Idk. Doesn’t really matter. It copied my string, except I probably want to copy the \0 too! This is probably the issue because I am no longer correctly terminating the string (for this exercise).

I add movsb below the jne .move and all is golden:

Expected ‘Work, please.’ Was ‘work’. Combined letters are: work
Expected ‘Horse, please.’ Was ‘horse’. Combined letters are: horse

I can now rewind the memory address and call front_door_password:

mov rdi, r8
call front_door_password

Expected ‘Work, please.’ Was ‘Work’. Combined letters are: work
Expected ‘Horse, please.’ Was ‘Horse’. Combined letters are: horse

Sweet. Now need to add the final piece of info.

I think what I can do is store ", please.\0" as an array in memory .odata, and then copy it at the end of the string. Quickly counting the size: 10 bytes including the final terminator.

section .odata
     politeness db `, please.`, 0
1 Like

First I store the string’s first position in rax, because I want to use stosb to copy from rax to rdi.

mov rax, [rel politeness]

I actually expected lea rax, [rel politeness] to work, but apparently it needs the data, not the memory address?

I’ll try to write 8 bytes first:

stosq
mov byte [rdi], 0

This gives me the expected result:

Expected ‘Work, please.’ Was ‘Work, please’. Combined letters are: work

I can decide to moving 2 more bytes (10 total)

stosq
stosw

Expected ‘Work, please.’ Was 'Work, please, '. Combined letters are: work

Ah. It copied 8 bytes and then again the same 2 bytes at the start. I mean, it didn’t say in the instructions that it would also increase rax by 1, but I suppose I hoped it did. Instead, I can try to use movs now.

mov rsi, [rel politeness]
movsq                ; move from rsi to rdi (8 bytes)
movsw                ; move from rsi to rdi (2 bytes)

Unfortunately no:

make: *** [Makefile:29: all] Segmentation fault (core dumped)

Hmmm. Maybe this needs to be lea rsi [rel politness] which is what I expected for stos. Let’s try.

Yeah. Okay.

I supposed I could also try to set rcx to the string length and use rep, and then move movsb instead. Let me try.

mov rcx, 10          ; prepare to move 10 characters (bytes)
movsb                ; move from rsi to rdi (1 byte)
rep                  ; repeat

Expected ‘Work, please.’ Was ‘Work,\x7F’. Combined letters are: work

This doesn’t work and I don’t understand why. I think this only moves one byte (and then it’s not null terminated, so the artifact is there). Is it not an instruction? Did prefix mean, just dump it in front? If so, that wasn’t completely clear and an example could help. Lemme try:

mov rcx, 10          ; prepare to move 10 characters (bytes)
repmovsb             ; move from rsi to rdi (1 byte * 10)

That yields a new error:

poetry_club_door_policy.asm:97: error: label alone on a line without a colon might be in error [-w+error=label-orphan]
make: *** [Makefile:35: poetry_club_door_policy.o] Error 1

Eh. Do I need to add a space? Aka is this an instruction that takes an instruction as an operand? That would be a first.

mov rcx, 10          ; prepare to move 10 characters (bytes)
rep movsb            ; move from rsi to rdi (1 byte * 10)

Passes all the tests again! I suppose I’ll leave it at that.

Oh final note. Yeah I can use that array length track $ - to get the length 10 and also store that. Maybe something for next time!

2 Likes

Thank you very much, @SleeplessByte !

I’ve already made a PR with a big rework of the concept in order to make things clearer. I’ll add an example of rep usage.

One thing is that you’re using section .odata, but it should be section .rodata (rodata = Read-Only DATA).

A section .odata works because you can create sections with non-conventional names (any name you want). But then, if you don’t specify the properties of this section, it is just a blob of raw data with the default settings (the same as section .data).

EDIT: The PR was just merged.

1 Like

Ah! Swear that’s what I read the first time around and didn’t understand the naming convention, but Read-Only makes more sense XD

2 Likes

12x better!

1 Like