Processing Large Data Streams in Ruby with Enumerator::Lazy
Why lazy enumeration matters
When a script has to chew through gigabytes of log lines or a CSV that does not fit in RAM, the naïve approach — reading the whole file into an array — quickly blows the memory budget. Ruby’s Enumerator::Lazy lets you build a pipeline that pulls one element at a time, transforms it, and passes it downstream, keeping the footprint tiny. I reach for it whenever the data source is a stream rather than a collection.
A real‑world log‑parsing scenario
Imagine a nightly job that scans an application log for requests that exceeded a latency threshold, extracts the request id, and writes a compact CSV for the analytics team. The log file can be several hundred megabytes, and the job runs on a modest CI runner with 512 MiB of RAM.
# frozen_string_literal: true
require 'csv'
THRESHOLD_MS = 200
INPUT_PATH = 'logs/app.log'
OUTPUT_PATH = 'reports/slow_requests.csv'
# Each line looks like:
# 2024-03-15T12:34:56.789Z [INFO] request_id=abc123 latency_ms=312
LINE_REGEX = /
^(?\S+)\s+
\[\w+\]\s+
request_id=(?\w+)\s+
latency_ms=(?\d+)
/x
def parse_line(line)
match = LINE_REGEX.match(line)
return nil unless match
{ req_id: match[:req_id], latency: match[:latency].to_i }
end
CSV.open(OUTPUT_PATH, 'w') do |csv|
csv << %w[request_id latency_ms]
File.foreach(INPUT_PATH).lazy # 1️⃣ lazy enumerator over lines
.map { |line| parse_line(line) } # 2️⃣ turn each line into a hash or nil
.compact # 3️⃣ drop lines that didn’t match
.select { |h| h[:latency] > THRESHOLD_MS } # 4️⃣ keep only the slow ones
.each { |h| csv << [h[:req_id], h[:latency]] }
end
The chain reads a single line, parses it, discards the noise, filters by latency, and writes the result — all without ever materialising the full log in memory. The lazy call on the enumerator returned by File.foreach is the only thing that changes the execution model; the rest of the pipeline looks exactly like a regular eager chain.
Putting it together
Why does this work? Enumerator::Lazy implements each by pulling the next value from the source only when the downstream consumer asks for it. Methods like map, select, and compact return new lazy enumerators that remember the transformation but defer execution. The final each (or to_a if you really need an array) triggers the pull‑based loop.
In the example above, CSV.open yields a writer object, and the block’s each drives the pipeline. Each iteration reads one line, runs the regex, decides whether to keep it, and immediately hands the row to the CSV writer. The GC sees only a handful of short‑lived objects at any moment.
Gotchas and tips
- Don’t call
to_aearly. That forces eager evaluation and defeats the purpose. - Side‑effects inside lazy steps are fine as long as they are idempotent per element (e.g., writing a row). Avoid mutating shared state.
- Combine with
Enumerator::Yielderfor custom sources — useful when you stream from an HTTP response or a database cursor. - Benchmark. For very small files the overhead of lazy wrappers can be slower than a simple
each_lineloop. Profile before you commit.
Lazy enumeration is not a silver bullet, but for any “process a huge stream once” task it turns a memory‑hungry script into a well‑behaved citizen.
Next time you see a File.readlines followed by a chain of map/select, ask yourself whether the data fits comfortably in RAM. If the answer is “maybe not”, wrap the source in .lazy and keep the pipeline identical. Your future self — and the ops team — will thank you.