Writing

Engineering notes.

Small notes on code, production systems, and lessons from building and maintaining software.

From practice

Two short reflections.

How I make code easier to follow and production behavior easier to reconstruct.
  1. 01

    Make the main method tell the story

    When an operation becomes complex, I want to understand its flow before I understand every implementation detail.

    I tend to organize multi-step business logic around a small use-case or orchestration object with one clear public entry point. The top-level method should read like a summary of the operation. Detailed work stays in focused methods; when one step develops its own dependencies, rules, and failure modes, I pull it into a separate component.

    Earlier in my career, I often stored every intermediate result on self. That made the control flow look clean, but it could hide which step produced the data consumed by the next one. These days I usually return and pass ordinary values explicitly. If a workflow genuinely has a lot of shared state, I would rather introduce a named execution or request context than keep adding implicit self attributes.

    For a simple transformation, a plain function is usually better. I reach for this structure when an operation has enough steps that the flow itself becomes something worth making explicit.

    A top-level operation with visible control and data flow
    class OrderCreator:
        def create(self, request):
            user = self._load_user(request)
            order = self._build_order(request, user)
            price = self._calculate_price(order)
            payment = self._charge(user, price)
            return self._finish(order, payment)
    The goal isn't to create more classes. It's to make complex code readable from the top down.
  2. 02

    Make every request traceable

    When something fails in production, I want to reconstruct what actually happened before relying on reproduction alone.

    For a long time, one of the simplest techniques that worked well for me was assigning each incoming request a correlation ID. If the caller already supplied one, I preserved it. Calls to downstream services carried the same ID in their headers, and every log entry included it. Once those logs reached Elasticsearch, one Kibana query could reconstruct a request across functions and services.

    I used to pass ctx_id explicitly through much of the application. It was clear and difficult to lose, but observability gradually leaked into business function signatures. Today I prefer the same idea with structured logging, request-scoped context, automatic propagation, and standard trace and span IDs where they fit. In Python, a request context or contextvars can let the logger attach that information without every function accepting another parameter.

    A correlation ID tells me which events belong together. Spans can also show who called whom and where time was spent. A small service may only need the first part; a distributed system may benefit from the full trace. The principle stayed the same; the implementation became more standardized.

    One request reconstructed across services
    POST /orders  trace_id=abc
    ├─ orders   event=create_order
    ├─ pricing  event=calculate_price
    └─ payment  event=charge_payment
       └─ ERROR timeout
    
    Kibana: trace.id = abc
    When production behaves unexpectedly, the system should leave enough evidence to reconstruct what happened.