Why lazy enumeration matters

When a script has to chew through gigabytes of log lines or a CSV that does not fit in RAM, the naïve approach — reading the whole file into an array — quickly blows the memory budget. Ruby’s Enumerator::Lazy lets you build a pipeline that pulls one element at a time, transforms it, and passes it downstream, keeping the footprint tiny. I reach for it whenever the data source is a stream rather than a collection.

A real‑world log‑parsing scenario

Imagine a nightly job that scans an application log for requests that exceeded a latency threshold, extracts the request id, and writes a compact CSV for the analytics team. The log file can be several hundred megabytes, and the job runs on a modest CI runner with 512 MiB of RAM.

# frozen_string_literal: true

require 'csv'

THRESHOLD_MS = 200
INPUT_PATH   = 'logs/app.log'
OUTPUT_PATH  = 'reports/slow_requests.csv'

# Each line looks like:
# 2024-03-15T12:34:56.789Z [INFO] request_id=abc123 latency_ms=312
LINE_REGEX = /
  ^(?\S+)\s+
  \[\w+\]\s+
  request_id=(?\w+)\s+
  latency_ms=(?\d+)
/x

def parse_line(line)
  match = LINE_REGEX.match(line)
  return nil unless match
  { req_id: match[:req_id], latency: match[:latency].to_i }
end

CSV.open(OUTPUT_PATH, 'w') do |csv|
  csv << %w[request_id latency_ms]
  File.foreach(INPUT_PATH).lazy          # 1️⃣ lazy enumerator over lines
    .map { |line| parse_line(line) }    # 2️⃣ turn each line into a hash or nil
    .compact                            # 3️⃣ drop lines that didn’t match
    .select { |h| h[:latency] > THRESHOLD_MS } # 4️⃣ keep only the slow ones
    .each { |h| csv << [h[:req_id], h[:latency]] }
end

The chain reads a single line, parses it, discards the noise, filters by latency, and writes the result — all without ever materialising the full log in memory. The lazy call on the enumerator returned by File.foreach is the only thing that changes the execution model; the rest of the pipeline looks exactly like a regular eager chain.

Putting it together

Why does this work? Enumerator::Lazy implements each by pulling the next value from the source only when the downstream consumer asks for it. Methods like map, select, and compact return new lazy enumerators that remember the transformation but defer execution. The final each (or to_a if you really need an array) triggers the pull‑based loop.

In the example above, CSV.open yields a writer object, and the block’s each drives the pipeline. Each iteration reads one line, runs the regex, decides whether to keep it, and immediately hands the row to the CSV writer. The GC sees only a handful of short‑lived objects at any moment.

Gotchas and tips

  • Don’t call to_a early. That forces eager evaluation and defeats the purpose.
  • Side‑effects inside lazy steps are fine as long as they are idempotent per element (e.g., writing a row). Avoid mutating shared state.
  • Combine with Enumerator::Yielder for custom sources — useful when you stream from an HTTP response or a database cursor.
  • Benchmark. For very small files the overhead of lazy wrappers can be slower than a simple each_line loop. Profile before you commit.

Lazy enumeration is not a silver bullet, but for any “process a huge stream once” task it turns a memory‑hungry script into a well‑behaved citizen.

Next time you see a File.readlines followed by a chain of map/select, ask yourself whether the data fits comfortably in RAM. If the answer is “maybe not”, wrap the source in .lazy and keep the pipeline identical. Your future self — and the ops team — will thank you.