When diving into low-level code or reading foundational books like Computer Systems: A Programmer's Perspective (CS:APP), you will frequently encounter assembly code generated by the GNU Compiler Collection (GCC). Sometimes, the output feels counter-intuitive.

Consider the following C function operating entirely on 64-bit signed integers (long in x86-64 Linux):

long arith(long x, long y, long z) {
    long t1 = x ^ y;
    long t2 = z * 48;
    long t3 = t1 & 0x0F0F0F0F;
    long t4 = t2 - t3;
    return t4;
}

When compiling with gcc -Og -S on x86-64, GCC produces the following assembly instructions:

    xorq    %rsi, %rdi
    leaq    (%rdx,%rdx,2), %rax
    salq    $4, %rax
    andl    $252645135, %edi
    subq    %rdi, %rax
    ret

Notice the fourth instruction: instead of using andq $0x0F0F0F0F, %rdi, the compiler emits andl $252645135, %edi. Why does GCC switch to the 32-bit register (%edi) and mnemonic (andl) when the variable t3 is a 64-bit long? Isn't it discarding the upper 32 bits?

The Key Rule: Implicit Zero-Extension in x86-64

To understand GCC's reasoning, you must understand a critical architectural design choice of the AMD64 (x86-64) architecture:

Any instruction that writes to a 32-bit register automatically clears (zero-extends) the upper 32 bits of the corresponding 64-bit register.

For example, executing an operation targeting %edi automatically sets bits 32–63 of %rdi to zero. You do not need a subsequent instruction to clear the top half of the register.

Does This Affect the Result for Large 64-Bit Numbers?

In short: no, the mathematical result is 100% identical.

The bitmask in our C code is 0x0F0F0F0F. When treated as an unsigned 64-bit value, it is equivalent to:

0x000000000F0F0F0F

If you perform a 64-bit bitwise AND between any arbitrary 64-bit integer and 0x000000000F0F0F0F, the upper 32 bits of the result will always evaluate to zero because x & 0 == 0.

Because the 32-bit instruction andl applies the mask to the lower 32 bits and automatically zeroes out the upper 32 bits of %rdi, the end state of the full 64-bit register %rdi is identical to what a 64-bit andq would produce.

Why GCC Prefers 32-Bit Instructions (Code Density)

Compilers are obsessed with instruction size and cache efficiency. Using 32-bit instructions instead of 64-bit instructions saves precious bytes in the resulting machine code:

  • No REX Prefix: 64-bit operand sizes in x86-64 generally require a 1-byte REX.W prefix (e.g., byte 0x48). 32-bit instructions do not require this prefix.
  • Smaller Immediate Operands: While both andl and andq can accept 32-bit sign-extended immediates, using the 32-bit variant saves encoding space whenever the top half of the immediate is zero.

Smaller instruction sizes mean better instruction cache (I-cache) utilization and faster decoding cycles by the CPU frontend.

Why Did Hexadecimal Change to Decimal?

You may also wonder why 0x0F0F0F0F became $252645135 in assembly.

This is simply a human-readable formatting convention of GCC’s assembly emitter. To the CPU and the assembler, numbers are just bit sequences:

  • Hexadecimal: 0x0F0F0F0F
  • Decimal: 252645135
  • Binary: 00001111000011110000111100001111

The GNU Assembler (gas) accepts either format. GCC frequently prints positive numerical constants as decimal values regardless of how you typed them in your C source code.

Summary

GCC emitted andl ..., %edi instead of andq ..., %rdi because:

  1. The bitmask fits comfortably in 32 bits with zeroes in the upper half.
  2. x86-64 automatically zero-extends 32-bit register writes into 64-bit registers.
  3. The 32-bit instruction requires fewer bytes of machine code, optimizing instruction cache usage.