Why Accumulation Matters

In day‑to‑day Ruby work I often find myself needing to walk through a collection and build up a result hash or array. The naïve approach is to declare an empty variable outside the loop, mutate it inside, and then return it after the iteration finishes. This works, but it mixes the responsibility of iteration with the responsibility of state management, making the code harder to reason about and test.

When the accumulator is a mutable object that lives outside the block, any future refactor that changes the loop structure can accidentally leave stale state behind. I’ve seen bugs where a developer added an early break or next and forgot to reset the accumulator, causing incorrect results in subsequent calls.

The Problem with Mutable Accumulators

Consider a typical snippet that tallies word frequencies:

def word_frequencies(words)
  freq = {}
  words.each do |w|
    freq[w] = (freq[w] || 0) + 1
  end
  freq
end

The method is short, but the freq variable is mutable and lives in the method scope. If I later decide to use each_with_index or add a guard clause, I have to remember to keep the mutation logic in sync. The mental overhead grows as the method becomes more complex.

What I really want is a way to express the accumulation as a pure transformation: give me an initial accumulator, yield each element and the current accumulator, and return the updated accumulator at the end. Ruby’s standard library already provides exactly that.

Enter each_with_object

Enumerable#each_with_object is designed for this pattern. It takes an initial object and passes it to the block together with each element. The block’s return value becomes the accumulator for the next iteration. When the enumeration finishes, the final accumulator is returned.

Here is the same word‑frequency example rewritten with each_with_object:

def word_frequencies(words)
  words.each_with_object(Hash.new(0)) do |w, hash|
    hash[w] += 1
  end
end

The initial accumulator is Hash.new(0), which automatically treats missing keys as zero. Inside the block we simply increment the count; we don’t need to fetch the old value or assign it back. The block’s return value is ignored—each_with_object uses the mutated hash directly for the next round.

Notice how the method now reads like a sentence: “Take the words, and for each word, with an empty hash that defaults to zero, increment the hash entry for that word.” The intent is clear, and there is no external mutable state to track.

Real‑World Example: Parsing a Sales CSV

Let’s look at a scenario I encounter regularly: summarising sales data from a CSV file. Each row contains a product ID, a region, and the amount sold. I need a nested hash that shows total sales per region, per product.

Using a mutable accumulator would look like this:

require 'csv'
def sales_summary(path)
  summary = {}
  CSV.foreach(path, headers: true) do |row|
    region = row['region']
    product = row['product_id']
    amount = row['amount'].to_f
    summary[region] ||= {}
    summary[region][product] ||= 0.0
    summary[region][product] += amount
  end
  summary
end

The code works, but the nested ||= assignments clutter the core logic. If I later need to filter rows by date, I have to be careful not to break the initialization lines.

With each_with_object the initialization can be encapsulated in a helper that returns a zero‑filled nested hash:

require 'csv'
def sales_summary(path)
  CSV.foreach(path, headers: true).each_with_object(Hash.new { |h, k| h[k] = Hash.new(0.0) }) do |row, summary|
    region = row['region']
    product = row['product_id']
    amount = row['amount'].to_f
    summary[region][product] += amount
  end
end

The outer Hash.new block creates a new Hash.new(0.0) for any unseen region, eliminating the need for ||=. Inside the iteration we only focus on the business rule: add the amount to the appropriate bucket.

This version is easier to test because the method is a pure function of the CSV rows; there are no hidden variables that could retain state between calls.

When to Reach for Other Tools

each_with_object shines when the accumulator is a single object that gets mutated. If you need to collect multiple independent results, consider map followed by flatten or reduce (aka inject) when you want to return a completely new object each iteration without mutation.

For example, building an array of transformed values is clearer with map:

def squared(numbers)
  numbers.map { |n| n * n }
end

But if you need to build a hash where keys depend on earlier values, each_with_object is often the most readable choice.

I keep a small snippet in my personal ~/.irbrc that defines a helper method hh returning Hash.new { |h, k| h[k] = Hash.new(0) }. It saves me a few keystrokes whenever I need a two‑level counter.

Adopting each_with_object has made my Ruby code feel more declarative and less prone to accidental state bugs. Give it a try in your next enumeration‑heavy method; you’ll likely find the intention of the code stands out more clearly, and future maintainers will thank you.