Why memoization matters

When a function repeatedly receives the same arguments, recomputing the result wastes CPU cycles and can hammer external services. In a recent project I built a service that called a third‑party pricing API for every order line. The API throttles at 100 requests per second, and many orders shared identical product IDs. Adding a tiny cache eliminated the bottleneck without any architectural changes.

Enter functools.lru_cache

Python’s standard library ships a least‑recently‑used (LRU) cache decorator that turns any pure function into a memoized version. It handles cache eviction, thread safety, and cache statistics out of the box.

Basic usage

from functools import lru_cache
import requests

@lru_cache(maxsize=256)
def fetch_price(product_id: str) -> float:
    """Retrieve the current price for a product."""
    resp = requests.get(f"https://api.example.com/price/{product_id}", timeout=5)
    resp.raise_for_status()
    return resp.json()["price"]

The decorator stores up to 256 distinct product IDs. Subsequent calls with the same ID return the cached float instantly.

Controlling cache size and lifetime

In production you often want a time‑based expiry rather than a pure LRU policy. A common pattern is to wrap the cached function and invalidate manually:

from functools import lru_cache
import time

class TimedCache:
    def __init__(self, ttl_seconds: int = 300):
        self.ttl = ttl_seconds
        self._cache = {}
        self._timestamps = {}

    def __call__(self, func):
        @lru_cache(maxsize=None)
        def wrapper(*args, **kwargs):
            key = (args, frozenset(kwargs.items()))
            now = time.time()
            if key in self._timestamps and now - self._timestamps[key] < self.ttl:
                return self._cache[key]
            result = func(*args, **kwargs)
            self._cache[key] = result
            self._timestamps[key] = now
            return result
        return wrapper

@TimedCache(ttl_seconds=60)
def fetch_price(product_id: str) -> float:
    ...

This gives you a 60‑second freshness window while still benefiting from LRU eviction when the cache grows.

Inspecting cache effectiveness

The decorated function exposes a cache_info() method. Logging it periodically reveals hit‑rate trends:

import logging
import time

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

def log_cache_stats():
    while True:
        info = fetch_price.cache_info()
        logger.info(
            "Cache stats – hits: %d, misses: %d, size: %d",
            info.hits, info.misses, info.currsize
        )
        time.sleep(60)

If the hit ratio stays low, you may be caching the wrong function or the maxsize is too small.

When not to use it

  • Mutable arguments – lists or dicts break hashing; convert to tuples or use a custom key function.
  • Side effects – functions that write to a DB, send emails, or modify global state must not be cached.
  • Highly dynamic data – if the underlying value changes every call, caching adds stale‑data risk.

Real‑world impact

After adding the timed cache to the pricing service, the API call volume dropped from ~8,000 requests/minute to under 1,200, well below the throttle limit. Latency for repeated product lookups fell from 120 ms to <2 ms. The change was a single decorator plus a tiny helper class — no new infrastructure, no config files.

Takeaway: functools.lru_cache is a zero‑dependency, battle‑tested tool for memoizing pure functions. Pair it with a lightweight TTL wrapper when you need freshness guarantees, and monitor cache_info() to verify the hit rate.