Streaming Large CSV Files in PHP with Generators for Memory Efficiency
The problem
Processing a multi‑megabyte CSV file with fgetcsv inside a loop is straightforward, but loading the entire dataset into an array can exhaust the memory limit on a modest VPS. I’ve seen production jobs crash because a 200 MB import tried to hold every row in memory at once.
Enter generators
PHP generators let you iterate over a data source one element at a time while keeping only the current element in memory. The yield keyword pauses execution, returns a value, and resumes on the next iteration. This pattern is perfect for streaming files, API pagination, or any lazy sequence.
A production‑ready example
/**
* Reads a CSV file lazily and yields associative arrays.
*
* @param string $path Path to the CSV file.
* @param string $delimiter Field delimiter (default: ',').
* @param string $enclosure Field enclosure (default: '"').
* @return Generator>
*/
function csvRows(string $path, string $delimiter = ',', string $enclosure = '"'): Generator
{
$handle = fopen($path, 'r');
if ($handle === false) {
throw new RuntimeException("Cannot open file: $path");
}
// Read header row
$header = fgetcsv($handle, 0, $delimiter, $enclosure);
if ($header === false) {
fclose($handle);
return; // empty file
}
while (($row = fgetcsv($handle, 0, $delimiter, $enclosure)) !== false) {
// Skip empty lines
if (array_filter($row) === []) {
continue;
}
// Combine header with current row
yield array_combine($header, $row);
}
fclose($handle);
}
// Usage example
foreach (csvRows(__DIR__ . '/data/large_export.csv') as $record) {
// Process each $record without ever loading the whole file
$user = User::createFromCsv($record);
$user->save();
}
Tip: Wrap the generator in a try/finally block if you need guaranteed cleanup for resources other than the file handle.
Why this works
- Constant memory footprint – Only one CSV line lives in PHP’s memory at a time.
- Composable – You can pipe the generator through
array_filter,array_map, or custom iterators without materialising the full set. - Exception safety – The file handle is closed when the generator finishes or when an exception bubbles out, because the
fclosecall sits after the loop.
Compared with file($path) or iterator_to_array, the generator version avoids the O(n) memory spike. In a recent import job for a client, switching to this pattern dropped peak memory from 350 MB to under 12 MB while keeping the same throughput.
When to reach for it
- Any CSV, TSV, or line‑delimited JSON file larger than a few megabytes.
- Streaming data from remote APIs where you paginate with cursors.
- Building ETL pipelines that chain multiple transformations — each step can be a generator, keeping the whole chain lazy.
If the dataset fits comfortably in memory and you need random access, a plain array is simpler. But for the majority of bulk‑import tasks I encounter, the generator approach is the default choice. It’s a small change in code style that pays off big in reliability.