The problem

Processing a multi‑megabyte CSV file with fgetcsv inside a loop is straightforward, but loading the entire dataset into an array can exhaust the memory limit on a modest VPS. I’ve seen production jobs crash because a 200 MB import tried to hold every row in memory at once.

Enter generators

PHP generators let you iterate over a data source one element at a time while keeping only the current element in memory. The yield keyword pauses execution, returns a value, and resumes on the next iteration. This pattern is perfect for streaming files, API pagination, or any lazy sequence.

A production‑ready example


/**
 * Reads a CSV file lazily and yields associative arrays.
 *
 * @param string $path Path to the CSV file.
 * @param string $delimiter Field delimiter (default: ',').
 * @param string $enclosure Field enclosure (default: '"').
 * @return Generator>
 */
function csvRows(string $path, string $delimiter = ',', string $enclosure = '"'): Generator
{
    $handle = fopen($path, 'r');
    if ($handle === false) {
        throw new RuntimeException("Cannot open file: $path");
    }

    // Read header row
    $header = fgetcsv($handle, 0, $delimiter, $enclosure);
    if ($header === false) {
        fclose($handle);
        return; // empty file
    }

    while (($row = fgetcsv($handle, 0, $delimiter, $enclosure)) !== false) {
        // Skip empty lines
        if (array_filter($row) === []) {
            continue;
        }
        // Combine header with current row
        yield array_combine($header, $row);
    }

    fclose($handle);
}

// Usage example
foreach (csvRows(__DIR__ . '/data/large_export.csv') as $record) {
    // Process each $record without ever loading the whole file
    $user = User::createFromCsv($record);
    $user->save();
}
Tip: Wrap the generator in a try/finally block if you need guaranteed cleanup for resources other than the file handle.

Why this works

  • Constant memory footprint – Only one CSV line lives in PHP’s memory at a time.
  • Composable – You can pipe the generator through array_filter, array_map, or custom iterators without materialising the full set.
  • Exception safety – The file handle is closed when the generator finishes or when an exception bubbles out, because the fclose call sits after the loop.

Compared with file($path) or iterator_to_array, the generator version avoids the O(n) memory spike. In a recent import job for a client, switching to this pattern dropped peak memory from 350 MB to under 12 MB while keeping the same throughput.

When to reach for it

  1. Any CSV, TSV, or line‑delimited JSON file larger than a few megabytes.
  2. Streaming data from remote APIs where you paginate with cursors.
  3. Building ETL pipelines that chain multiple transformations — each step can be a generator, keeping the whole chain lazy.

If the dataset fits comfortably in memory and you need random access, a plain array is simpler. But for the majority of bulk‑import tasks I encounter, the generator approach is the default choice. It’s a small change in code style that pays off big in reliability.